Ying Shen 0005

dblp:01/8558-5 · DBLP profile ↗
← Back
49ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0002-2966-7955ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images
abstract
Recent advances in generative AI have accelerated the production of ultra-high-resolution visual content. However, traditional image formats face significant limitations in efficient compression and real-time decoding, which restricts their applicability on end-user devices. Inspired by 3D Gaussian Splatting, 2D Gaussian image models have achieved notable progress in enhancing image representation efficiency and quality. Nevertheless, existing methods struggle to balance compression ratios and reconstruction fidelity in ultra-high-resolution scenarios. To address these challenges, we propose SmartSplat, a highly adaptive and feature-aware GS-based image compression framework that effectively supports arbitrary image resolutions and compression ratios. By leveraging image-aware features such as gradients and color variances, SmartSplat introduces a Gradient-Color Guided Variational Sampling strategy alongside an Exclusion-based Uniform Sampling scheme, significantly improving the non-overlapping coverage of Gaussian primitives in pixel space. Additionally, a Scale-Adaptive Gaussian Color Sampling method is proposed to enhance the initialization of Gaussian color attributes across scales. Through joint optimization of spatial layout, scale, and color initialization, SmartSplat can efficiently capture both local structures and global textures of images using a limited number of Gaussians, achieving superior reconstruction quality under high compression ratios. Extensive experiments on DIV8K and a newly created 16K dataset demonstrate that SmartSplat significantly outperforms state-of-the-art methods at comparable compression ratios and surpasses their compression limits, exhibiting strong scalability and practical applicability. This framework can effectively alleviate the storage and transmission burdens of ultra-high-resolution images, providing a robust foundation for future high-efficiency visual content processing.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
AAAI4
2026 Why Do Emotions Change? Appraisal-Guided Reasoning for Emotion-Cause Triplet Extraction in Conversations
abstract
Multimodal Emotion-Cause Triplet Extraction in Conversations (MECTEC) is fundamental for fine-grained affect understanding, yet it remains challenging in multi-turn, multispeaker settings.Existing methods often make locally plausible predictions but struggle to maintain conversation-level consistency under within-speaker emotion shifts and core events.To address this, we propose ECFlow, a unified framework that combines appraisal-guided structured generation with graph-structured reinforcement learning.ECFlow operationalizes cognitive appraisal theory into a controllable intermediate reasoning trace and constructs UMECS, a unified supervision dataset with cognitively grounded traces.It then lifts predicted and gold triplets into an Emotion-Cause Flow Graph and optimizes verifiable, structure-aware rewards for emotion-shift coherence and core-event consistency, together with task-oriented triplet rewards.Experiments on public MECTEC benchmarks show that ECFlow consistently outperforms strong baselines, achieving state-of-the-art triplet extraction and improved structure-aware metrics on emotion shifts and core events.
Ying Shen 0005, Lin Zhang 0014
ACL (1)2
2026 Decision-Invariant Sim-to-Real Vision-and-Language Navigation with Pseudo-Panoramic Observations
abstract
Following natural language instructions to complete navigation tasks is a crucial capability for real-world embodied robots. In vision-and-language navigation (VLN), agents typically assume access to complete, on-the-fly environmental observations and rely on them for decision making. However, most sim-to-real VLN approaches approximate the privileged complete panoramic sensing in simulation with monocular sensors, leading to significant semantic loss and incomplete perception. In this work, we present DIP2, a sim-to-real VLN framework that reduces the discrepancy between assumed and realizable observations by introducing a unified sensing and mapping representation shared across simulated and real-world domains. Specifically, dense pseudo-panoramic observations are synthesized from a self-assembled surround-view camera system to recover rich semantics comparable to simulation, while a structure-only metric map unifies simulated RGB-D inputs with real-world LiDAR observations. Furthermore, to promote sim-to-real decision invariance, DIP2 employs a learning-free local navigable waypoint prediction strategy via radial expansion, applied consistently in both domains. The predicted waypoints and dense semantic observations are integrated into a global topological map, allowing existing agents trained in simulation to be directly deployed in the real world. Extensive experiments in simulated and real environments demonstrate that DIP2 substantially improves sim-to-real navigation performance. Source code will be published at https://github.com/zheng19845/DIP2.
Yuanyu Zheng, Xumin Shen, Yunda Sun, Ying Shen 0005, Lin Zhang 0014
ICMR4
2026 ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
abstract
Large Language Models (LLMs) have shown strong potential as conversational agents. Yet, their effectiveness remains limited by deficiencies in robust long-term memory, particularly in complex, long-term web-based services such as online emotional support. However, existing long-term dialogue benchmarks primarily focus on static and explicit fact retrieval, failing to evaluate agents in critical scenarios where user information is dispersed, implicit, and continuously evolving. To address this gap, we introduce ES-MemEval, a comprehensive benchmark that systematically evaluates five core memory capabilities: information extraction, temporal reasoning, conflict detection, abstention, and user modeling, in long-term emotional support settings, covering question answering, summarization, and dialogue generation tasks. To support the benchmark, we also propose EvoEmo, a multi-session dataset for personalized long-term emotional support that captures fragmented, implicit user disclosures and evolving user states. Extensive experiments on open-source long-context, commercial, and retrieval-augmented (RAG) LLMs show that explicit long-term memory is essential for reducing hallucinations and enabling effective personalization. At the same time, RAG improves factual consistency but struggles with temporal dynamics and evolving user states. These findings highlight both the potential and limitations of current paradigms and motivate more robust integration of memory and retrieval for long-term personalized dialogue systems.
Jiaqi Lu 0004, Ying Shen 0005, Lin Zhang 0014
WWW3
2026 CaneSpeaker: An LLM-Assisted Speaker for Generating Human-Like Navigation Instructions
abstract
Navigation instruction generation aims to address data scarcity in Vision-and-Language Navigation (VLN) by generating navigation instructions for unannotated routes from data sources like simulators or online data. However, existing methods usually suffer from high reliance on panoramic views, poor cross-task generalization ability, and limited availability of training data. To address these challenges, we propose a novel speaker, CaneSpeaker, to generate human-like instructions from front-facing images for a variety of VLN tasks. First, to mitigate the limited amount of speaker training data, we propose an Large Language Model (LLM)-based instruction augmentation method, LLM-IA, that utilizes an off-the-shelf LLM to create augmented instructions for training by distilling and reformulating existing instructions. This method allows us to collect an instruction-augmented dataset with human-level accuracy for speaker training, namely Rx2R. Second, to eliminate the dependency on panoramic views, we propose a novel Vision-Language Model (VLM)-based speaker architecture, VL-Sp. By leveraging the advanced reasoning capabilities of a pre-trained VLM, CaneSpeaker can effectively generate high-quality instructions directly from front-facing images without relying on panoramic views. Also, the prompt-based characteristic of the VLM allows us to devise a unified input representation to enable the processing of multiple VLN tasks, thus further addressing the problem of data scarcity by combining multiple datasets from different VLN tasks. Finally, we utilize CaneSpeaker to synthesize a large-scale augmented dataset, CANE, from unannotated routes in the Matterport3D Simulator. Comprehensive experiments demonstrate that CaneSpeaker generates precise instructions with diverse expressions across various VLN tasks, and the VLN agent trained on our datasets obviously outperforms its counterparts. The source codes and datasets are available at https://github.com/zheng19845/CaneSpeaker .
Yuanyu Zheng, Lin Zhang 0014, Yunda Sun, Ying Shen 0005, Shengjie Zhao 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Representing Sounds as Neural Amplitude Fields: A Benchmark of Coordinate-MLPs and a Fourier Kolmogorov-Arnold Framework
abstract
Although Coordinate-MLP-based implicit neural representations have excelled in representing radiance fields, 3D shapes, and images, their application to audio signals remains underexplored. To fill this gap, we investigate existing implicit neural representations, from which we extract 3 types of positional encoding and 16 commonly used activation functions. Through combinatorial design, we establish the first benchmark for Coordinate-MLPs in audio signal representations. Our benchmark reveals that Coordinate-MLPs require complex hyperparameter tuning and frequency-dependent initialization, limiting their robustness. To address these issues, we propose Fourier-ASR, a novel framework based on the Fourier series theorem and the Kolmogorov-Arnold representation theorem. Fourier-ASR introduces Fourier Kolmogorov-Arnold Networks (Fourier-KAN), which leverage periodicity and strong nonlinearity to represent audio signals, eliminating the need for additional positional encoding. Furthermore, a Frequency-adaptive Learning Strategy (FaLS) is proposed to enhance the convergence of Fourier-KAN by capturing high-frequency components and preventing overfitting of low-frequency signals. Extensive experiments conducted on natural speech and music datasets reveal that: (1) well-designed positional encoding and activation functions in Coordinate-MLPs can effectively improve audio representation quality; and (2) Fourier-ASR can robustly represent complex audio signals without extensive hyperparameter tuning. Looking ahead, the continuity and infinite resolution of implicit audio representations make our research highly promising for tasks such as audio compression, synthesis, and generation.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
AAAI6
2025 Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-Sources
abstract
Audio-visual navigation has received considerable attention in recent years. However, the majority of related investigations have focused on single sound-source scenarios. Studies in this field for multiple sound-source scenarios remain underexplored due to the limitations of two aspects. First, the existing audio-visual navigation dataset only has limited audio samples, making it difficult to simulate diverse multiple sound-source environments. Second, existing navigation frameworks are mainly designed for single sound-source scenarios, thus their performance is severely reduced in multiple sound-source scenarios. In this work, we make an attempt to fill in these two research gaps to some extent. First, we establish a large-scale BEnchmark Dataset for Audio-Vsual Navigation, namely BeDAViN. This dataset consists of 2,258 audio samples with a total duration of 10.8 hours, which is more than 33 times longer than the existing audio dataset employed in the audio-visual navigation task. Second, we propose a new Embodied Navigation framework for MUltiple Sound-Sources Scenarios called ENMuS3. There are mainly two essential components in ENMuS3, the sound event descriptor and the multi-scale scene memory transformer. The former component equips the agent with the ability to extract spatial and semantic features of the target sound-source among multiple sound-sources, while the latter provides the ability to track the target object effectively in noisy environments. Experimental results on our BeDAViN show that ENMuS3 strongly outperforms its counterparts with a significant improvement in success rates across diverse scenarios.
Zhanbo Shi, Lin Zhang 0014, Linfei Li, Ying Shen 0005
AAAI4
2025 WSGS: A Speech-Driven Zero-Shot System for 6D Robotic Arm Grasping
abstract
In the robotic vision industry, recent years have witnessed a growing interest in object detection and pose estimation for precise robotic arm grasping. In fact, various unfavorable factors, such as the size limitations of robotic grippers, the diversity and potential complexity of object shapes and poses, and the cluttered environment, make robotic arm grasping based on 6D object poses much harder than it seems. In this paper, to solve these issues to some extent, we proposed a speech-driven zero-shot system for robotic arm grasping, called WSGS (Whisper-SAM6D Grasping System). It enables speech-driven human-interactive grasping with the Franka Emika robotic arm by performing instance segmentation and pose estimation based on speech instructions. Specifically, WSGS accurately recognizes the 6D pose of an unknown object and adapts to its size to find the most suitable position for grasping. Comprehensive experiments on our real-world scenarios demonstrate that WSGS can produce high-accuracy instance segmentation and pose estimation results, achieving adaptive robotic arm grasping of unknown objects based on speech commands.
Yitong Ge, Lin Zhang 0014, Yang Chen 0037, Ying Shen 0005
ICME4
2024 GS3LAM: Gaussian Semantic Splatting SLAM
abstract
Recently, the multi-modal fusion of RGB, depth, and semantics has shown great potential in the domain of dense Simultaneous Localization and Mapping (SLAM), as known as dense semantic SLAM. Yet a prerequisite for generating consistent and continuous semantic maps is the availability of dense, efficient, and scalable scene representations. To date, existing semantic SLAM systems based on explicit scene representations (points/meshes/surfels) are limited by their resolutions and inabilities to predict unknown areas, thus failing to generate dense maps. Contrarily, a few implicit scene representations (Neural Radiance Fields) to deal with these problems rely on time-consuming ray tracing-based volume rendering technique, which cannot meet the real-time rendering requirements of SLAM. Fortunately, the Gaussian Splatting scene representation has recently emerged, which inherits the efficiency and scalability of point/surfel representations while smoothly represents geometric structures in a continuous manner, showing promise in addressing the aforementioned challenges. To this end, we propose GS3LAM, a Gaussian Semantic Splatting SLAM framework, which takes multimodal data as input and can render consistent, continuous dense semantic maps in real-time. To fuse multimodal data, GS3LAM models the scene as a Semantic Gaussian Field (SG-Field), and jointly optimizes camera poses and the field by establishing error constraints between observed and predicted data. Furthermore, a Depth-adaptive Scale Regularization (DSR) scheme is proposed to tackle the problem of misalignment between scale-invariant Gaussians and geometric surfaces within the SG-Field. To mitigate the forgetting phenomenon, we propose an effective Random Sampling-based Keyframe Mapping (RSKM) strategy, which exhibits notable superiority over local covisibility optimization strategies commonly utilized in 3DGS-based SLAM systems. Extensive experiments conducted on the benchmark datasets reveal that compared with state-of-the-art competitors, GS3 LAM demonstrates increased tracking robustness, superior real-time rendering quality, and enhanced semantic reconstruction precision. To make the results reproducible, the source code is available at https://github.com/lif314/GS3LAM.
Linfei Li, Lin Zhang 0014, Zhong Wang 0009, Ying Shen 0005
ACM Multimedia4
2024 MPEG: A Multi-Perspective Enhanced Graph Attention Network for Causal Emotion Entailment in Conversations
abstract
Emotion causes constitute a pivotal component in the comprehension of emotional conversations. Recently, a new task named Causal Emotion Entailment (CEE) has been proposed to identify the causal utterances for the target emotional utterance in a conversation. Although researchers have achieved some progress in solving this problem, they failed to adequately incorporate speaker characteristics and overlooked the effects of temporal relations in conversation structures. To fill such a research gap to some extent, we propose a novel causal emotion entailment framework, namely MPEG (Multi-Perspective Enhanced Graph attention network). The training of MPEG consists of three stages. Firstly, we utilize a speaker-aware pre-trained model and two attention mechanisms to obtain the utterance representations that incorporate local contexts as well as the speaker and emotional information. Then, these representations are fed into a graph attention network to model the conversation structures and emotional dynamics from both local and global perspectives. Finally, a fully-connected network is implemented to predict the relationships between emotional utterances and causal utterances. Experimental results show that MPEG achieves state-of-the-art performance. The source code is available athttps://github.com/slptongji/MPEG.
Ying Shen 0005, Xuri Chen, Lin Zhang 0014, Shengjie Zhao 0001
IEEE Trans. Affect. Comput.2
2024 TriKF: Triple-Perspective Knowledge Fusion Network for Empathetic Question Generation
abstract
Questioning is one of the essential tactics for demonstrating empathy in social dialogues. Effective questioning can guide individuals to express their experiences, feelings, and thoughts, aiming to establish emotional connections and deepen interpersonal understanding. However, how to generate empathetic questions in emotional support conversations remains an unresolved issue. To fill this research gap to some extent, we propose an empathetic question generation (QG) framework called triple-perspective knowledge fusion (TriKF), which incorporates external knowledge from the perspectives of events, cognition, and affection to comprehensively understand the dialogue context. Specifically, this framework acquires commonsense knowledge from these three perspectives and integrates them into the dialogue context to enrich the contextual information. To the best of our knowledge, this is the first method proposed for empathetic QG. Additionally, we construct an empathetic question dataset, namely EQ-EMAC. This dataset comprises 4213 dialogues with single user inputs and multiple empathetic question responses, which can be utilized to assess the effectiveness and generalization capability of empathetic QG models. Experimental results have demonstrated the effectiveness of TriKF on the task of empathetic QG compared with seven baseline models.
Ying Shen 0005, Xuri Chen, Lin Zhang 0014, Shengjie Zhao 0001
IEEE Trans. Comput. Soc. Syst.2
2023 Extrinsic Self-Calibration of the Surround-View System: A Weakly Supervised Approach
abstract
An SVS usually consists of four wide-angle fisheye cameras mounted around the vehicle to sense the surrounding environment. From the images synchronously captured by all cameras, a top-down surround-view can be synthesized, on the premise that both intrinsics and extrinsics of the cameras have been calibrated. At present, the intrinsic calibration approach is relatively well-developed and can be pipelined, while the extrinsic calibration is still immature. On one hand, the existing manual calibration schemes are usually reliable, but need to be conducted by professionals in specific sites, which is undoubtedly cumbersome. On the other hand, the majority of the existing self- calibration schemes are based on low-level features and their stability and robustness are usually unsatisfactory. As far as we know, an effective extrinsic self-calibration scheme designed specially for the SVS is still lacking. To fill such a research gap to some extent, we propose a novel self-calibration scheme which follows a weakly supervised framework, namely WESNet (Weakly-supervised Extrinsic Self-calibration Network). The training of WESNet consists of two stages. First, we utilize the corners in a few calibration site images as the weak supervision to roughly optimize the network by minimizing the geometric loss. Then, after the convergency in the first stage, we additionally introduce a self-supervised photometric loss term that can be constructed by the photometric information from natural images for further fine-tuning. Besides, to support training, we totally collected 19,078 groups of synchronously captured fisheye images under various environmental conditions. To our knowledge, thus far this is the largest surround-view dataset containing original fisheye images. By means of learning prior knowledge from the training data, WESNet takes the original fisheye images synchronously collected as the input, and directly yields extrinsics end-to-end with little labor cost. Its efficiency and efficacy have been corroborated by extensive experiments conducted on our collected dataset. To make our results reproducible, source code and the collected dataset have been released at https://cslinzhang.github.io/WESNet/WESNet.html.
Yang Chen 0037, Lin Zhang 0014, Ying Shen 0005, Brian Nlong Zhao, Yicong Zhou
IEEE Trans. Multim.3
2023 D-LIOM: Tightly-Coupled Direct LiDAR-Inertial Odometry and Mapping
abstract
Simultaneous localization and mapping via LiDAR-Inertial fusion is a crucial technology in many automation-related applications. Recently, a number of approaches based on geometric features have evolved, yielding impressive results via tightly-coupled estimation. This sort of feature-based techniques, however, are inextricably linked to the scanning mechanism of the LiDAR, relying on stable feature detection, and thus are difficult to adapt to multi-LiDAR systems. A few “direct” solutions, on the other hand, register the raw point cloud with the built probability map, which is more computationally efficient and easy to be extended. But, the existing direct approaches are all loosely-coupled, lacking correction of the IMU biases, and thus only work well in 2D cases. To this end, we present D-LIOM, a tightly-coupled Direct LiDAR-Inertial Odometry and Mapping framework. In D-LIOM, a scan is directly registered to a probability submap, and the LiDAR odometry, the IMU pre-integration, and the gravity constraint are integrated to build a local factor graph in the submap's time window, allowing the system to perform real-time high-precision pose estimation. Furthermore, to eliminate accumulated errors in time, we detect loops and adjust the sparse pose graph based on mutual matching of projected 2D submaps, allowing D-LIOM to run stably in large-scale scenes. In addition, to improve its flexibility to varied sensor combinations, D-LIOM supports multi-LiDAR inputs and facilitates the initialization with a common 6-axis IMU. Extensive experiments demonstrate that D-LIOM largely outperforms the existing state-of-the-art counterparts in mapping effect and localization accuracy as well as with high time efficiency. Lastly, to ensure that our results are entirely reproducible, all necessary data and codes are made open-source available. One introduction video can also be found on the online website.
Zhong Wang 0009, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
IEEE Trans. Multim.3
2023 SLAM for Indoor Parking: A Comprehensive Benchmark Dataset and a Tightly Coupled Semantic Framework
abstract
For the task of autonomous indoor parking, various Visual-Inertial Simultaneous Localization And Mapping (SLAM) systems are expected to achieve comparable results with the benefit of complementary effects of visual cameras and the Inertial Measurement Units. To compare these competing SLAM systems, it is necessary to have publicly available datasets, offering an objective way to demonstrate the pros/cons of each SLAM system. However, the availability of such high-quality datasets is surprisingly limited due to the profound challenge of the groundtruth trajectory acquisition in the Global Positioning Satellite denied indoor parking environments. In this article, we establish BeVIS, a large-scale Be nchmark dataset with V isual (front-view), I nertial and S urround-view sensors for evaluating the performance of SLAM systems developed for autonomous indoor parking, which is the first of its kind where both the raw data and the groundtruth trajectories are available. In BeVIS, the groundtruth trajectories are obtained by tracking artificial landmarks scattered in the indoor parking environments, whose coordinates are recorded in a surveying manner with a high-precision Electronic Total Station. Moreover, the groundtruth trajectories are comprehensively evaluated in terms of two respects, the reprojection error and the pose volatility, respectively. Apart from BeVIS, we propose a novel tightly coupled semantic SLAM framework, namely VIS SLAM -2, leveraging V isual (front-view), I nertial, and S urround-view sensor modalities, specially for the task of autonomous indoor parking. It is the first work attempting to provide a general form to model various semantic objects on the ground. Experiments on BeVIS demonstrate the effectiveness of the proposed VIS SLAM -2. Our benchmark dataset BeVIS is publicly available at https://shaoxuan92.github.io/BeVIS .
Xuan Shao, Ying Shen 0005, Lin Zhang 0014, Shengjie Zhao 0001, Dandan Zhu 0001, Yicong Zhou
ACM Trans. Multim. Comput. Commun. Appl.2
2022 Towards Controllable and Physical Interpretable Underwater Scene Simulation
abstract
The realistic simulation of underwater scenes has important significance for many researches related to underwater vision, such as underwater image restoration, underwater moving object monitoring, etc. To date, however, the existing underwater scene simulation pipelines are either too complicated due to the continuous spectra and camera parameters involved, or difficult to control since the empirically controlled distance based fog effect is usually used by them. In this paper, we try to fill in this research gap by proposing an Underwater Scene Simulation approach, namely USSim, which especially focuses on the influence of ocean water. In USSim, Jerlov water type and depth are regarded as main variables to control the simulation effects. In addition, the spectra of the incident light is decomposed into three primary components and their attenuations are modeled separately, and finally the simulated scene is generated via the hybrid underwater imaging model proposed by us. USSim greatly reduces the computational complexity and enables the fog effect to be controlled by variables with explicit physical meanings. The controllability, physical interpretability and simulation effects of our USSim under different conditions have been verified by extensive experiments. To make our results reproducible, the source code is made online available at https://cslinzhang.github.io/USSim/.
Kaixin Chen 0003, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
ICASSP3
2022 Chunkfusion: A Learning-Based RGB-D 3D Reconstruction Framework Via Chunk-Wise Integration
abstract
Recent years have witnessed a growing interest in online RGB-D 3D reconstruction. On the premise of ensuring the reconstruction accuracy with noisy depth scans, making the system scalable to various environments is still challenging. In this paper, we devote our efforts to try to fill in this research gap by proposing a scalable and robust RGB-D 3D reconstruction framework, namely Chunk-Fusion. In ChunkFusion, sparse voxel management is exploited to improve the scalability of online reconstruction. Besides, a chunk-wise TSDF (truncated signed distance function) fusion network is designed to perform a robust integration of the noisy depth measurements on the sparsely allocated voxel chunks. The proposed chunk-wise TSDF integration scheme can accurately restore surfaces with superior visual consistency from noisy depth maps and can guarantee the scalability of online reconstruction simultaneously, making our reconstruction framework widely applicable to scenes with various scales and depth scans with strong noises and outliers. The outstanding scalability and efficacy of our ChunkFusion have been corroborated by extensive experiments. To make our results reproducible, the source code is made online available at https://cslinzhang.github.io/ChunkFusion/.
Chaozheng Guo, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
ICASSP3
2022 Automatic Depression Detection: an Emotional Audio-Textual Corpus and A Gru/Bilstm-Based Model
abstract
Depression is a global mental health problem, the worst case of which can lead to suicide. An automatic depression detection system provides great help in facilitating depression self-assessment and improving diagnostic accuracy. In this work, we propose a novel depression detection approach utilizing speech characteristics and linguistic contents from participants’ interviews. In addition, we establish an Emotional Audio-Textual Depression Corpus (EATD-Corpus) which contains audios and extracted transcripts of responses from depressed and non-depressed volunteers. To the best of our knowledge, EATD-Corpus is the first and only public depression dataset that contains audio and text data in Chinese. Evaluated on two depression datasets, the proposed method achieves the state-of-the-art performances. The outperforming results demonstrate the effectiveness and generalization ability of the proposed method. The source code and EATD-Corpus are available at https://github.com/speechandlanguageprocessing/ICASSP2022-Depression.
Ying Shen 0005, Huiyu Yang
ICASSP1
2022 SiD-WaveFlow: A Low-Resource Vocoder Independent of Prior Knowledge
Ying Shen 0005, Dongqing Wang, Lin Zhang 0014
INTERSPEECH2
2022 LVI-ExC: A Target-free LiDAR-Visual-Inertial Extrinsic Calibration Framework
abstract
Recently, the multi-modal fusion with 3D LiDAR, camera, and IMU has shown great potential in applications of automation-related fields. Yet a prerequisite for a successful fusion is that the geometric relationships among the sensors are accurately determined, which is called an extrinsic calibration problem. To date, the existing target-based approaches to deal with this problem rely on sophisticated calibration objects (sites) and well-trained operators, which is time-consuming and inflexible in practical applications. Contrarily, a few target-free methods can overcome these shortcomings, while they only focus on the calibrations of two types of the sensors. Although it is possible to obtain LiDAR-visual-inertial extrinsics by chained calibrations, problems such as cumbersome operations, large cumulative errors, and weak geometric consistency still exist. To this end, we propose LVI-ExC, an integrated LiDAR-Visual-Inertial Extrinsic Calibration framework, which takes natural multi-modal data as input and yields sensor-to-sensor extrinsics end-to-end without any auxiliary object (site) or manual assistance. To fuse multi-modal data, we formulate the LiDAR-visual-inertial extrinsic calibration as a continuous-time simultaneous localization and mapping problem, in which the extrinsics, trajectories, time differences, and map points are jointly estimated by establishing sensor-to-sensor and sensor-to-trajectory constraints. Extensive experiments show that LVI-ExC can produce precise results. With LVI-ExC's outputs, the LiDAR-visual reprojection results and the reconstructed environment map are all highly consistent with the actual natural scenes, demonstrating LVI-ExC's outstanding performance. To ensure that our results are fully reproducible, all the relevant data and codes have been released publicly at https://cslinzhang.github.io/LVI-ExC/.
Zhong Wang 0009, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
ACM Multimedia3
2022 MOFISSLAM: A Multi-Object Semantic SLAM System With Front-View, Inertial, and Surround-View Sensors for Indoor Parking
abstract
The semantic SLAM (Simultaneous Localization And Mapping) system is a crucial module for autonomous indoor parking. Visual cameras (monocular/binocular) and IMU (Inertial Measurement Unit) constitute the basic configuration to build such a system. The performance of existing SLAM systems typically deteriorates in the presence of dynamically movable objects or objects with little texture. By contrast, semantic objects on the ground embody the most salient and stable features in the indoor parking environment. Due to their inabilities to perceive such features on the ground, existing SLAM systems are prone to tracking inconsistency during navigation. In this paper, we present MOFISSLAM, a novel tightly-coupled${M}$ulti-${O}$bject semantic SLAM system integrating${F}$ront-view,${I}$nertial, and${S}$urround-view sensors for autonomous indoor parking. The proposed system moves beyond existing semantic SLAM systems by complementing the sensor configuration with a surround-view system capturing images from a top-down viewpoint. In MOFISSLAM, apart from low-level visual features and inertial motion data, typical semantic objects (parking-slots, parking-slot IDs and speed bumps) detected in surround-views are also incorporated in optimization, forming robust surround-view constraints. Specifically, each surround-view feature imposes a surround-view constraint that can be split into a contact term and a registration term. The former pre-defines the position of each individual surround-view feature subject to whether it has semantic contact with other surround-view features. Three contact modes, defined ascomplementary,adjacentandcoincident, are identified to guarantee a unified form of all contact terms. The latter further constrains by registering each surround-view observation and its position in the world coordinate system. In parallel, to objectively evaluate SLAM studies for autonomous indoor parking, a large-scale dataset with groundtruth trajectories is collected, which is the first of its kind. Its groundtruth trajectories, commonly unavailable, are obtained by tracking artificial features scattered in the indoor parking environment, whose 3D coordinates are measured with an ETS (Electronic Total Station). The collected dataset has been made publicly available athttps://shaoxuan92.github.io/MOFIS.
Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Yicong Zhou
IEEE Trans. Circuits Syst. Video Technol.4
2021 ROECS: A Robust Semi-direct Pipeline Towards Online Extrinsics Correction of the Surround-view System
abstract
Generally, a surround-view system (SVS), which is an indispensable component of advanced driving assistant systems (ADAS), consists of four to six wide-angle fisheye cameras. As long as both intrinsics and extrinsics of all cameras have been calibrated, a top-down surround-view with the real scale can be synthesized at runtime from fisheye images captured by these cameras. However, when the vehicle is driving on the road, relative poses between cameras in the SVS may change from the initial calibrated states due to bumps or collisions. In case that extrinsics' representations are not adjusted accordingly, on the surround-view, obvious geometric misalignment will appear. Currently, the researches on correcting the extrinsics of the SVS in an online manner are quite sporadic, and a mature and robust pipeline is still lacking. As an attempt to fill this research gap to some extent, in this work, we present a novel extrinsics correction pipeline designed specially for the SVS, namely ROECS (Robust Online Extrinsics Correction of the Surround-view system). Specifically, a "refined bi-camera error" model is firstly designed. Then, by minimizing the overall "bi-camera error" within a sparse and semi-direct framework, the SVS's extrinsics can be iteratively optimized and become accurate eventually. Besides, an innovative three-step pixel selection strategy is also proposed. The superior robustness and the generalization capability of ROECS are validated by both quantitative and qualitative experimental results. To make the results reproducible, the collected data and the source code have been released at https://cslinzhang.github.io/ROECS/.
Tianjun Zhang, Brian Nlong Zhao, Ying Shen 0005, Xuan Shao, Lin Zhang 0014, Yicong Zhou
ACM Multimedia3
2021 RefineDNet: A Weakly Supervised Refinement Framework for Single Image Dehazing
abstract
Haze-free images are the prerequisites of many vision systems and algorithms, and thus single image dehazing is of paramount importance in computer vision. In this field, prior-based methods have achieved initial success. However, they often introduce annoying artifacts to outputs because their priors can hardly fit all situations. By contrast, learning-based methods can generate more natural results. Nonetheless, due to the lack of paired foggy and clear outdoor images of the same scenes as training samples, their haze removal abilities are limited. In this work, we attempt to merge the merits of prior-based and learning-based approaches by dividing the dehazing task into two sub-tasks, i.e., visibility restoration and realness improvement. Specifically, we propose a two-stage weakly supervised dehazing framework, RefineDNet. In the first stage, RefineDNet adopts the dark channel prior to restore visibility. Then, in the second stage, it refines preliminary dehazing results of the first stage to improve realness via adversarial learning with unpaired foggy and clear images. To get more qualified results, we also propose an effective perceptual fusion strategy to blend different dehazing outputs. Extensive experiments corroborate that RefineDNet with the perceptual fusion has an outstanding haze removal capability and can also produce visually pleasing results. Even implemented with basic backbone networks, RefineDNet can outperform supervised dehazing approaches as well as other state-of-the-art methods on indoor and outdoor datasets. To make our results reproducible, relevant code and data are available at https://github.com/xiaofeng94/RefineDNet-for-dehazing.
Shiyu Zhao 0001, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou
IEEE Trans. Image Process.3
2020 A Study Of Parking-Slot Detection With The Aid Of Pixel-Level Domain Adaptation
abstract
The self-parking system is an important component of self-driving vehicles. Such a system needs to detect and locate the parking-slots from surround-view images, and then guide the vehicle to the designated parking-slot. In the real world, the appearances and environmental conditions of parking-slots can be rich and varied. Thus, to train the parking-slot detection model, it is necessary to collect and label a huge quantity of surround-view images covering as many real cases as possible. Such a process is cumbersome and costly, and will be repeated whenever encountering an unseen parking condition that is quite different from the ones covered by existing training set. To this end, in this paper we propose an extensible pipeline, namely FakePS, to assist parking-slot detection model training by making use of synthetic data. Specifically, with FakePS, we can first build various simulated parking scenes and collect labeled surround-view images automatically. Besides, we resort to pixel-level domain adaptation strategies to enhance the realism of the synthetic images using unlabeled real images while preserving their label information. The efficacy of FakePS has been corroborated by experimental results.
Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou
ICME3
2020 Oecs: Towards Online Extrinsics Correction For The Surround-View System
abstract
A typical surround-view system consists of four fisheye cameras. By performing an offline calibration that determines both the intrinsics and extrinsics of the system, surround-view images can be synthesized at runtime. However, poses of calibrated cameras sometimes may change. In such a case, if cameras' extrinsics are not updated accordingly, observable geometric misalignment will appear in surround-views. Most existing solutions to this problem resort to re-calibration, which is quite cumbersome. Thus, how to correct cameras' extrinsics in an online manner without using re-calibration is still an open issue. In this paper, we attempt to propose a novel solution to this problem and the proposed solution is referred to as “Online Extrinsics Correction for the Surround-view system OECS for short. We first design a Bi-Camera error model, measuring the photometric discrepancy between two corresponding pixels on images captured by two adjacent cameras. Then, by minimizing the system's overall BiCamera error, cameras' extrinsics can be optimized and the optimization is conducted within a sparse direct framework. The efficacy and efficiency of OECS are validated by experiments. Data and source code used in this work are publicly available at https://z619850002.github.io/OECage/.
Tianjun Zhang, Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou
ICME3
2020 Zero-Shot Restoration of Underexposed Images via Robust Retinex Decomposition
abstract
Underexposed images often suffer from serious quality degradation such as poor visibility and latent noise in the dark. Most previous methods for underexposed images restoration ignore the noise and amplify it during stretching contrast. We predict the noise explicitly to achieve the goal of denoising while restoring the underexposed image. Specifically, a novel three-branch convolution neural network, namely RRDNet (short for Robust Retinex Decomposition Network), is proposed to decompose the input image into three components, illumination, reflectance and noise. As an image-specific network, RRDNet doesn't need any prior image examples or prior training. Instead, the weights of RRDNet will be updated by a zero-shot scheme of iteratively minimizing a specially designed loss function. Such a loss function is devised to evaluate the current decomposition of the test image and guide noise estimation. Experiments demonstrate that RRDNet can achieve robust correction with overall naturalness and pleasing visual quality. To make the results reproducible, the source code has been made publicly available at https://aaaaangel.github.io/RRDNet-Homepage.
Lin Zhang 0014, Ying Shen 0005, Yong Ma 0005, Shengjie Zhao 0001, Yicong Zhou
ICME3
2020 A Tightly-coupled Semantic SLAM System with Visual, Inertial and Surround-view Sensors for Autonomous Indoor Parking
abstract
The semantic SLAM (simultaneous localization and mapping) system is an indispensable module for autonomous indoor parking. Monocular and binocular visual cameras constitute the basic configuration to build such a system. Features used in existing SLAM systems are often dynamically movable, blurred and repetitively textured. By contrast, semantic features on the ground are more stable and consistent in the indoor parking environment. Due to their inabilities to perceive salient features on the ground, existing SLAM systems are prone to tracking loss during navigation. Therefore, a surround-view camera system capturing images from a top-down viewpoint is necessarily called for. To this end, this paper proposes a novel tightly-coupled semantic SLAM system by integrating Visual, Inertial, and Surround-view sensors, VIS SLAM for short, for autonomous indoor parking. In VIS SLAM, apart from low-level visual features and IMU (inertial measurement unit) motion data, parking-slots in surround-view images are also detected and geometrically associated, forming semantic constraints. Specifically, each parking-slot can impose a surround-view constraint that can be split into an adjacency term and a registration term. The former pre-defines the position of each individual parking-slot subject to whether it has an adjacent neighbor. The latter further constrains by registering between each observed parking-slot and its position in the world coordinate system. To validate the effectiveness and efficiency of VIS SLAM, a large-scale dataset composed of synchronous multi-sensor data collected from typical indoor parking sites is established, which is the first of its kind. The collected dataset has been made publicly available at https://cslinzhang.github.io/VISSLAM/.
Xuan Shao, Lin Zhang 0014, Tianjun Zhang, Ying Shen 0005, Hongyu Li 0001, Yicong Zhou
ACM Multimedia4
2020 Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and Baselines
abstract
On benchmark images, modern dehazing methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. Thus, it is imperative to adopt quantitative evaluation on a vast number of hazy images. However, existing quantitative evaluation schemes are not convincing due to a lack of appropriate datasets and poor correlations between metrics and human perceptions. In this work, we attempt to address these issues, and we make two contributions. First, we establish two benchmark datasets, i.e., the BEnchmark Dataset for Dehazing Evaluation (BeDDE) and the EXtension of the BeDDE (exBeDDE), which had been lacking for a long period of time. The BeDDE is used to evaluate dehazing methods via full reference image quality assessment (FR-IQA) metrics. It provides hazy images, clear references, haze level labels, and manually labeled masks that indicate the regions of interest (ROIs) in image pairs. The exBeDDE is used to assess the performance of dehazing evaluation metrics. It provides extra dehazed images and subjective scores from people. To the best of our knowledge, the BeDDE is the first dehazing dataset whose image pairs were collected in natural outdoor scenes without any simulation. Second, we provide a new insight that dehazing involves two separate aspects, i.e., visibility restoration and realness restoration, which should be evaluated independently; thus, to characterize them, we establish two criteria, i.e., the visibility index (VI) and the realness index (RI), respectively. The effectiveness of the criteria is verified through extensive experiments. Furthermore, 14 representative dehazing methods are evaluated as baselines using our criteria on BeDDE. Our datasets and relevant code are available at https://github.com/xiaofeng94/BeDDE-for-defogging.
Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001
IEEE Trans. Image Process.4
2019 Seamless 3D Surround View with a Novel Burger Model
abstract
In recent years, the 3D surround view (3D-SV) system has become a hot research topic in the field of Advanced Driver Assistance Systems (ADAS). It can be used to form a stereoscopic view of the surrounding 3D environment by using 4 car-mounted surround cameras, and users can switch the viewpoints for virtual observation conveniently. However, there are still many problems in how to stitch calibrated images to the panoramic view and how to project the panorama to the surround view. In this paper, we introduce the graph cut algorithm and multi-band blending to the panorama stitching phase. In addition, we design a new hamburger-shaped 3D geometric model to be the carrier of the panorama for texture mapping. Our 3D-SV system can make drivers have an immersive visual experience. Experimental results show that the 3D-SV generated by our method is less distorted and looks more natural than the other competitors.
Lin Zhang 0014, Ying Shen 0005, Shengjie Zhao 0001
ICIP4
2019 DMPR-PS: A Novel Approach for Parking-Slot Detection Using Directional Marking-Point Regression
abstract
The self-parking system plays an important role in autonomous driving, and one of its critical issues is parking-slot detection. Previous studies in this field are mostly based on off-the-shelf models designed for universal purposes, which have various limitations in solving specific problems. In this paper, we propose a parking-slot detection method using directional marking-point regression, namely DMPR-PS. Instead of utilizing multiple off-the-shelf models, DMPR-PS uses a novel CNN-based model specially designed for directional marking-point regression. Given a surround-view image I, the model predicts position, shape and orientation of each marking-point on I. From marking-points, parking-slots on I could be easily inferred using geometric rules. DMPR-PS outperforms state-of-the-art competitors on the benchmark dataset with a precision rate of 99.42% and a recall rate of 99.37%, while achieving a real-time detection speed of 12ms per frame on Nvidia Titan Xp. To make the results reproducible, the source code is available at https://github.com/Teoge/DMPR-PS.
Lin Zhang 0014, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang
ICME3
2019 Revisit Surround-view Camera System Calibration
abstract
The surround-view system is an essential component of an advanced driver assistance system especially when the vehicle runs in tight parking space or on a narrow road. To ensure successful maneuvering, a panoramic bird's-eye image with no blind spots is necessarily called for. Hence, a typical surround-view system consists of several cameras mounted around the vehicle capturing images from a top-down viewpoint, and an accurate extrinsic calibration for such system is prerequisite for providing a seamless surround-view image. To achieve this goal, this paper presents a novel extrinsic calibration pipeline which is both easy-to-use and reliable to operate on multiple cameras. Instead of taking the vehicle to a fixed position in a specific calibration site, a single chessboard is the only demand. We adopt a novel refinement procedure that jointly optimizes camera poses in a closed-loop manner. The effectiveness and efficiency of the proposed pipeline to calibrate a surround-view camera system has been corroborated by experiments.
Xuan Shao, Xiao Liu 0030, Lin Zhang 0014, Shengjie Zhao 0001, Ying Shen 0005, Yukai Yang
ICME5
2019 Pay By Showing Your Palm: A Study of Palmprint Verification on Mobile Platforms
abstract
With the fast development of smart mobile devices, mobile phones have gradually become an indispensable part of people's lives. Many biometric technologies based on mobile platforms have also developed rapidly, such as face verification and fingerprint recognition. However, the great potential of palmprint has been neglected. In this paper, we conducted a thorough study of palmprint verification on mobile devices for the first time. Firstly, we established an annotated, palmprint dataset named MPD, which was collected by multi-brands phones in two different sessions. As the largest dataset in this field, MPD contains 16,000 palm images from 200 subjects. Secondly, we built a DCNN-based palmprint verification system named DeepMPV for mobile platforms. The efficiency and performance of our system have been corroborated on our collected dataset. The labelled dataset and the source code are publicly available at https://cslinzhang.github.io/deepmpv/.
Lin Zhang 0014, Xiao Liu 0030, Shengjie Zhao 0001, Ying Shen 0005, Yukai Yang
ICME5
2019 Evaluation of Defogging: A Real-World Benchmark Dataset, A New Criterion and Baselines
abstract
Modern defogging methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. On the other hand, existing quantitative evaluation methods are also not convincing due to a lack of proper datasets. In this work, we attempt to address these issues and establish a long-term lacking benchmark dataset, namely BeDDE (BEnchmark Dataset for Defogging Evaluation), for evaluating the performance of defogging algorithms. To our knowledge, BeDDE is the first real-world dataset comprising foggy images with their registered clear counterparts. Using BeDDE, we set up a new criterion for evaluating defogging methods where VSI, a full reference image quality assessment metric, is calculated and averaged on registered ROIs of all image pairs. The evaluation results of the proposed criterion correlate well with human judgements. 10 state-of-the-art defogging methods are evaluated as baselines on BeDDE. BeDDE is available online.
Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang
ICME4
2019 Online Camera Pose Optimization for the Surround-view System
abstract
Surround-view system is an important information medium for drivers to monitor the driving environment. A typical surround-view system consists of four to six fish-eye cameras arranged around the vehicle. From these camera inputs, a top-down image of the ground around the vehicle, namely the surround-view image can be generated with well calibrated camera poses. Although existing surround-view system solutions can estimate camera poses accurately in off-line environment, how to correct the camera poses' change in online environment is still an open issue. In this paper, we propose a camera pose optimization method for surround-view system in online environment. Our method consists of two models: Ground Model and Ground-Camera Model, both of which correct the camera poses by minimizing photometric errors between ground projections of adjacent cameras. Experiments show that our method can effectively correct the geometric misalignment of the surround-view image caused by camera poses' change. Since our method is highly automated with low requirement of calibration site and manual operation, it has a wide range of applications and is convenient for the end-users. To make the results reproducible, the source code is publicly available at https://cslinzhang.github.io/CamPoseOpt/.
Xiao Liu 0030, Lin Zhang 0014, Ying Shen 0005, Shaoming Zhang, Shengjie Zhao 0001
ACM Multimedia3
2019 Zero-Shot Restoration of Back-lit Images Using Deep Internal Learning
abstract
How to restore back-lit images still remains a challenging task. State-of-the-art methods in this field are based on supervised learning and thus they are usually restricted to specific training data. In this paper, we propose a "zero-shot" scheme for back-lit image restoration, which exploits the power of deep learning, but does not rely on any prior image examples or prior training. Specifically, we train a small image-specific CNN, namely ExCNet (short for Exposure Correction Network) at test time, to estimate the "S-curve" that best fits the test back-lit image. Once the S-curve is estimated, the test image can be then restored straightforwardly. ExCNet can adapt itself to different settings per image. This makes our approach widely applicable to different shooting scenes and kinds of back-lighting conditions. Statistical studies performed on 1512 real back-lit images demonstrate that our approach can outperform the competitors by a large margin. To the best of our knowledge, our scheme is the first unsupervised CNN-based back-lit image restoration method. To make the results reproducible, the source code is available at https://cslinzhang.github.io/ExCNet/.
Lin Zhang 0014, Lijun Zhang 0005, Xiao Liu 0030, Ying Shen 0005, Shaoming Zhang, Shengjie Zhao 0001
ACM Multimedia4
2018 A CNN-Based Depth Estimation Approach with Multi-scale Sub-pixel Convolutions and a Smoothness Constraint
Shiyu Zhao 0001, Lin Zhang 0014, Ying Shen 0005, Yongning Zhu
ACCV (2)3
2018 Image Exposure Assessment: A Benchmark and a Deep Convolutional Neural Networks Based Model
abstract
In the camera equipment manufacturing industry, the exposure calibration is one of the basic steps for manufacturers to consider before launching their products to the market. To this end, a method that can objectively and automatically assess the exposure levels of images taken by the camera is highly desired. However, few studies have been conducted in this area. In this paper, we attempt to solve this issue to some extent and our contributions are twofold. Firstly, in order to facilitate the study of image exposure assessment, an Image Exposure Database$(IE_{ps}D)$is established. In this database, there are 15, 582 images with various exposure levels, and for each image there is an associated subjective exposure score which could reflect its perceptual exposure level. Secondly, we propose a novel highly accurate DCNN-based model, namely$IE_{ps}M$(Image Exposure Metric), to predict the exposure level of a given image.
Lijun Zhang 0005, Lin Zhang 0014, Xiao Liu 0030, Ying Shen 0005, Dongqing Wang
ICME4
2017 The Hasp Motif: A New Type of RNA Tertiary Interactions
Ying Shen 0005, Lin Zhang 0014
ICIC (2)1
2017 Vision-based parking-slot detection: A benchmark and a learning-based approach
abstract
Recent years have witnessed a growing interest in developing automatic parking systems in the field of intelligent vehicle. However, how to effectively and efficiently locating parking-slots using a vision-based system is still an unresolved issue. In this paper, we attempt to fill this research gap to some extent and our contributions are twofold. Firstly, to facilitate the study of vision-based parking-slot detection, a large-scale parking-slot image database is established. For each image in this database, the marking-points and parking-slots are carefully labelled. Such a database can serve as a benchmark to design and validate parking-slot detection algorithms. Secondly, a learning based parking-slot detection approach is proposed. With this approach, given a test image, the marking-points will be detected at first and then the valid parking-slots can be inferred. Its efficacy and efficiency have been corroborated on our database. The labeled database and the source codes are publicly available at http://sse.tongji.edu.cn/linzhang/ps/index.htm.
Linshen Li, Lin Zhang 0014, Xiyuan Li, Xiao Liu 0030, Ying Shen 0005
ICME5
2017 Image set classification based on synthetic examples and reverse training
Lin Zhang 0014, Qingjun Liang, Ying Shen 0005, Meng Yang 0001, Feng Liu 0013
Neurocomputing3
2017 Towards contactless palmprint recognition: A novel device, a new benchmark, and a collaborative representation based identification approach
Lin Zhang 0014, Lida Li, Anqi Yang, Ying Shen 0005, Meng Yang 0001
Pattern Recognit.4
2015 The λ-Turn: A New Structural Motif in Ribosomal RNA
Huizhu Ren, Ying Shen 0005, Lin Zhang 0014
ICIC (2)2
2015 RNA-binding residues prediction using structural features
abstract
BACKGROUND: RNA-protein complexes play an essential role in many biological processes. To explore potential functions of RNA-protein complexes, it's important to identify RNA-binding residues in proteins. RESULTS: In this work, we propose a set of new structural features for RNA-binding residue prediction. A set of template patches are first extracted from RNA-binding interfaces. To construct structural features for a residue, we compare its surrounding patches with each template patch and use the accumulated distances as its structural features. These new features provide sufficient structural information of surrounding surface of a residue and they can be used to measure the structural similarity between the surface surrounding two residues. The new structural features, together with other sequence features, are used to predict RNA-binding residues using ensemble learning technique. CONCLUSIONS: The experimental results reveal the effectiveness of the proposed structural features. In addition, the clustering results on template patches exhibit distinct structural patterns of RNA-binding sites, although the sequences of template patches in the same cluster are not conserved. We speculate that RNAs may have structure preferences when binding with proteins.
Huizhu Ren, Ying Shen 0005
BMC Bioinform.2
2015 3D Palmprint Identification Using Block-Wise Features and Collaborative Representation
abstract
Developing 3D palmprint recognition systems has recently begun to draw attention of researchers. Compared with its 2D counterpart, 3D palmprint has several unique merits. However, most of the existing 3D palmprint matching methods are designed for one-to-one verification and they are not efficient to cope with the one-to-many identification case. In this paper, we fill this gap by proposing a collaborative representation (CR) based framework with l1-norm or l2-norm regularizations for 3D palmprint identification. The effects of different regularization terms have been evaluated in experiments. To use the CR-based classification framework, one key issue is how to extract feature vectors. To this end, we propose a block-wise statistics based feature extraction scheme. We divide a 3D palmprint ROI into uniform blocks and extract a histogram of surface types from each block; histograms from all blocks are then concatenated to form a feature vector. Such feature vectors are highly discriminative and are robust to mere misalignment. Experiments demonstrate that the proposed CR-based framework with an l2-norm regularization term can achieve much better recognition accuracy than the other methods. More importantly, its computational complexity is extremely low, making it quite suitable for the large-scale identification application. Source codes are available at http://sse.tongji.edu.cn/linzhang/cr3dpalm/cr3dpalm.htm.
Lin Zhang 0014, Ying Shen 0005, Hongyu Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 VSI: A Visual Saliency-Induced Index for Perceptual Image Quality Assessment
abstract
Perceptual image quality assessment (IQA) aims to use computational models to measure the image quality in consistent with subjective evaluations. Visual saliency (VS) has been widely studied by psychologists, neurobiologists, and computer scientists during the last decade to investigate, which areas of an image will attract the most attention of the human visual system. Intuitively, VS is closely related to IQA in that suprathreshold distortions can largely affect VS maps of images. With this consideration, we propose a simple but very effective full reference IQA method using VS. In our proposed IQA model, the role of VS is twofold. First, VS is used as a feature when computing the local quality map of the distorted image. Second, when pooling the quality score, VS is employed as a weighting function to reflect the importance of a local region. The proposed IQA index is called visual saliency-based index (VSI). Several prominent computational VS models have been investigated in the context of IQA and the best one is chosen for VSI. Extensive experiments performed on four large-scale benchmark databases demonstrate that the proposed IQA index VSI works better in terms of the prediction accuracy than all state-of-the-art IQA indices we can find while maintaining a moderate computational complexity. The MATLAB source code of VSI and the evaluation results are publicly available online at http://sse.tongji.edu.cn/linzhang/IQA/VSI/VSI.htm.
Lin Zhang 0014, Ying Shen 0005, Hongyu Li 0001
IEEE Trans. Image Process.2
2013 Improving Classification Accuracy Using Gene Ontology Information
Ying Shen 0005, Lin Zhang 0014
ICIC (3)1
2012 Generalized Adjusted Rand Indices for cluster ensembles
Shaohong Zhang, Hau-San Wong, Ying Shen 0005
Pattern Recognit.3
2012 A New Unsupervised Feature Ranking Method for Gene Expression Data Based on Consensus Affinity
abstract
Feature selection is widely established as one of the fundamental computational techniques in mining microarray data. Due to the lack of categorized information in practice, unsupervised feature selection is more practically important but correspondingly more difficult. Motivated by the cluster ensemble techniques, which combine multiple clustering solutions into a consensus solution of higher accuracy and stability, recent efforts in unsupervised feature selection proposed to use these consensus solutions as oracles. However,these methods are dependent on both the particular cluster ensemble algorithm used and the knowledge of the true cluster number. These methods will be unsuitable when the true cluster number is not available, which is common in practice. In view of the above problems, a new unsupervised feature ranking method is proposed to evaluate the importance of the features based on consensus affinity. Different from previous works, our method compares the corresponding affinity of each feature between a pair of instances based on the consensus matrix of clustering solutions. As a result, our method alleviates the need to know the true number of clusters and the dependence on particular cluster ensemble approaches as in previous works. Experiments on real gene expression data sets demonstrate significant improvement of the feature ranking results when compared to several state-of-the-art techniques.
Shaohong Zhang, Hau-San Wong, Ying Shen 0005, Dongqing Xie
IEEE ACM Trans. Comput. Biol. Bioinform.3
2011 Feature-based 3D motif filtering for ribosomal RNA
abstract
MOTIVATION: RNA 3D motifs are recurrent substructures in an RNA subunit and are building blocks of the RNA architecture. They play an important role in binding proteins and consolidating RNA tertiary structures. RNA 3D motif searching consists of two steps: candidate generation and candidate filtering. We proposed a novel method, known as Feature-based RNA Motif Filtering (FRMF), for identifying motifs based on a set of moment invariants and the Earth Mover's Distance in the second step. RESULTS: A positive set of RNA motifs belonging to six characteristic types, with eight subtypes occurring in HM 50S, is compiled by us. The proposed method is validated on this representative set. FRMF successfully finds most of the positive fragments. Besides the proposed new method and the compiled positive set, we also recognize some new motifs, in particular a π-turn and some non-standard A-minor motifs are found. These newly discovered motifs provide more information about RNA structure conformation. AVAILABILITY: Matlab code can be downloaded from www.cs.cityu.edu.hk/~yingshen/FRMF.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ying Shen 0005, Hau-San Wong, Shaohong Zhang, Zhiwen Yu 0002
Bioinform.1
2010 A new method for measuring the semantic similarity on gene ontology
abstract
Semantic similarity defined on Gene Ontology (GO) aims to provide the functional relationship between different biological processes, molecular functions, or cellular components. In this paper, a novel method, namely the Shortest Path (SP) algorithm, for measuring the semantic similarity on GO is proposed based on both the GO structure information and the term's property. The proposed algorithm searches for the shortest path that connects two terms and uses the sum of weights on the shortest path to compute the semantic similarity for GO terms. A method for evaluating the nonlinear correlation between two variables is also introduced for validation. Extensive experiments conducted on two public gene expression datasets demonstrate the overall superiority of SP method over the other state-of-the-art methods evaluated.
Ying Shen 0005, Shaohong Zhang, Hau-San Wong
BIBM1