Siyuan Xiang

dblp:94/8355 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Knowledge representation and reasoning · 52% Robot manipulation · 30% 3D vision · 9%
Computer networks
2 papers
Content delivery and video streaming · 55% Wireless networking · 26% Physical-layer communications · 19%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning
1.022022
Self-supervised Spatial Reasoning on Multi-View Line Drawings · CVPR 2022
SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings · CVPR 2020
Robotics › Robot manipulation › manipulation skills
pouring
0.612022
Look and Listen: A Multi-Sensory Pouring Network and Dataset for Granular Media from Human Demonstrations · ICRA 2022
Content delivery and video streaming
adaptive video streaming
0.212014
A Real-Time Adaptive Algorithm for Video Streaming over Multiple Wireless Access Networks · IEEE J. Sel. Areas Commun. 2014
Wireless networking
link selection
0.212014
A Real-Time Adaptive Algorithm for Video Streaming over Multiple Wireless Access Networks · IEEE J. Sel. Areas Commun. 2014
Computer vision › 3D vision
camera pose estimation
0.212022
Self-supervised Spatial Reasoning on Multi-View Line Drawings · CVPR 2022
Computer vision › Segmentation and scene understanding
multisensory perception
0.212022
Look and Listen: A Multi-Sensory Pouring Network and Dataset for Granular Media from Human Demonstrations · ICRA 2022
Physical-layer communications
modulation
0.112010
Scalable Modulation for Scalable Wireless Videocast · INFOCOM 2010
Content delivery and video streaming › video coding
scalable video coding
0.112010
Scalable Modulation for Scalable Wireless Videocast · INFOCOM 2010
Content delivery and video streaming › video multicast
wireless video multicast
0.112010
Scalable Modulation for Scalable Wireless Videocast · INFOCOM 2010
Multimedia systems and quality of experience
quality of service
0.112014
A Real-Time Adaptive Algorithm for Video Streaming over Multiple Wireless Access Networks · IEEE J. Sel. Areas Commun. 2014
Physical-layer communications
channel coding
0.012010
Scalable Modulation for Scalable Wireless Videocast · INFOCOM 2010

Methods — techniques the papers use, named apart from their topics

self-supervised classification · 0.6neural network · 0.6contrastive learning · 0.6audio-visual fusion · 0.6resnet baseline · 0.4dataset generation · 0.4scalable video coding · 0.4markov decision process · 0.4best-action search · 0.4simulation · 0.1cross-layer optimization · 0.1
YearPublicationVenuePosition
2022 Self-supervised Spatial Reasoning on Multi-View Line Drawings
abstract
Spatial reasoning on multi-view line drawings by state-of-the-art supervised deep networks is recently shown with puzzling low performances on the SPARE3D dataset [14]. Based on the fact that self-supervised learning is helpful when a large number of data are available, we propose two self-supervised learning approaches to improve the baseline performance for view consistency reasoning and camera pose reasoning tasks on the SPARE3D dataset. For the first task, we use a self-supervised binary classification network to contrast the line drawing differences between various views of any two similar 3D objects, enabling the trained networks to effectively learn detail-sensitive yet view-invariant line drawing representations of 3D objects. For the second type of task, we propose a self-supervised multi-class classification framework to train a model to select the correct corresponding view from which a line drawing is rendered. Our method is even helpful for the downstream tasks with unseen camera poses. Experiments show that our method could significantly increase the baseline performance in SPARE3D, while some popular self-supervised learning methods cannot.
Siyuan Xiang, Anbang Yang, Yanfei Xue, Yaoqing Yang 0002, Chen Feng 0002
CVPR1
2022 Look and Listen: A Multi-Sensory Pouring Network and Dataset for Granular Media from Human Demonstrations
abstract
Humans have the ability to pour various media, both liquid and granular, to desired ends in various containers. We do this by using multiple senses simultaneously in a constant feedback loop to complete a pouring task. Combining multiple sensing modalities, similar to humans, could aid in robotic pouring control outside of a structured or industrial setting. We present a multi-sensory pouring dataset consisting of human pouring demonstrations of various granular media, coupled with two multi-sensory networks that estimate pouring rate and pouring average height. For both pouring metrics, a combined input of audio and visual data provides a lower median error than either the audio network or visual network. The multi-sensory network achieves a median error of 6.4 mm for average height estimation and 0.06 N/s for pouring rate estimation.
Alexis Burns, Siyuan Xiang, Dae-Won Lee, Lawrence D. Jackel, Shuran Song, Volkan Isler
ICRA2
2020 SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings
abstract
Spatial reasoning is an important component of human intelligence. We can imagine the shapes of 3D objects and reason about their spatial relations by merely looking at their three-view line drawings in 2D, with different levels of competence. Can deep networks be trained to perform spatial reasoning tasks? How can we measure their "spatial intelligence"? To answer these questions, we present the SPARE3D dataset. Based on cognitive science and psychometrics, SPARE3D contains three types of 2D-3D reasoning tasks on view consistency, camera pose, and shape generation, with increasing difficulty. We then design a method to automatically generate a large number of challenging questions with ground truth answers for each task. They are used to provide supervision for training our baseline models using state-of-the-art architectures like ResNet. Our experiments show that although convolutional networks have achieved superhuman performance in many visual learning tasks, their spatial reasoning performance in SPARE3D is almost equal to random guesses. We hope SPARE3D can stimulate new problem formulations and network designs for spatial reasoning to empower intelligent robots to operate effectively in the 3D world via 2D sensors.
Wenyu Han, Siyuan Xiang, Ruoyu Wang 0012, Chen Feng 0002
CVPR2
2015 Dynamic rate adaptation for adaptive video streaming in wireless networks
Siyuan Xiang, Min Xing, Lin Cai 0001, Jianping Pan 0001
Signal Process. Image Commun.1
2014 A Real-Time Adaptive Algorithm for Video Streaming over Multiple Wireless Access Networks
abstract
Video streaming is gaining popularity among mobile users. The latest mobile devices, such as smart phones and tablets, are equipped with multiple wireless network interfaces. How to efficiently and cost-effectively utilize multiple links to improve video streaming quality needs investigation. In order to maintain high video streaming quality while reducing the wireless service cost, in this paper, the optimal video streaming process with multiple links is formulated as a Markov Decision Process (MDP). The reward function is designed to consider the quality of service (QoS) requirements for video traffic, such as the startup latency, playback fluency, average playback quality, playback smoothness and wireless service cost. To solve the MDP in real time, we propose an adaptive, best-action search algorithm to obtain a sub-optimal solution. To evaluate the performance of the proposed adaptation algorithm, we implemented a testbed using the Android mobile phone and the Scalable Video Coding (SVC) codec. Experiment results demonstrate the feasibility and effectiveness of the proposed adaptation algorithm for mobile video streaming applications, which outperforms the existing state-of-the-art adaptation algorithms.
Min Xing, Siyuan Xiang, Lin Cai 0001
IEEE J. Sel. Areas Commun.2
2013 Transmission Control for Compressive Sensing Video over Wireless Channel
abstract
In this paper, we consider a wireless sensor node monitoring the environment and it is equipped with a compressive-sensing based, single-pixel image camera and other sensors such as temperature and humidity sensors. The wireless node needs to send the data out in a timely and energy efficient way. This transmission control problem is challenging in that we need to jointly consider perceived video quality, quality variation, power consumption and transmission delay requirements, and the wireless channel uncertainty. We address the above issues by first building a rate-distortion model for compressive sensing video. Then we formulate the deterministic and stochastic optimization problems and design the transmission control algorithm which jointly performs rate control, scheduling and power control. Extensive simulations have been conducted to demonstrate the effectiveness of the proposed transmission control algorithm.
Siyuan Xiang, Lin Cai 0001
IEEE Trans. Wirel. Commun.1
2012 Rate adaptation strategy for video streaming over multiple wireless access networks
abstract
Video streaming is gaining popularity among mobile users. The latest mobile devices, such as smart phones and tablets are equipped with multiple wireless network interfaces. How to efficiently and cost-effectively utilize multiple links to improve the video streaming quality needs to be investigated. In order to maintain high video streaming quality while reduce the wireless service cost, in this paper, the optimal video streaming process with multiple links is formulated as a Markov Decision Process (MDP). The reward function is designed to consider the quality of experience (QoE) requirements for video traffic, such as the interruption rate, average playback quality, playback smoothness and wireless service cost. Using dynamic programming, the MDP can be solved to obtain the optimal streaming policy. To evaluate the performance of the proposed multi-link rate adaptation (MLRA) algorithm, we implement a testbed using the Android mobile phone and the open-source X264 video codec. Experimental results demonstrate the feasibility and effectiveness of the proposed MLRA algorithm for mobile video streaming applications, which outperforms the existing state-of-the-art one.
Min Xing, Siyuan Xiang, Lin Cai 0001
GLOBECOM2
2012 Adaptive scalable video streaming in wireless networks
abstract
In this paper, we investigate the optimal streaming strategy for dynamic adaptive streaming over HTTP (DASH). Specifically, we focus on the rate adaptation algorithm for streaming scalable video (H.264/SVC) in wireless networks. We model the rate adaptation problem as a Markov Decision Process (MDP), aiming to find an optimal streaming strategy in terms of user-perceived quality of experience (QoE) such as playback interruption, average playback quality and playback smoothness. We then obtain the optimal MDP solution using dynamic programming. We further define a reward parameter in our proposed streaming strategy, which can be adjusted to make a good trade-off between the average playback quality and playback smoothness. We also use a simple testbed to validate our solution. Experiment results show the feasibility of the proposed solution and its advantage over the existing work.
Siyuan Xiang, Lin Cai 0001, Jianping Pan 0001
MMSys1
2011 Scalable Video Coding with Compressive Sensing for Wireless Videocast
abstract
Channel coding such as Reed-Solomon (RS) and convolutional codes has been widely used to protect video transmission in wireless networks. However, this type of channel coding can effectively correct error bits only if the error rate is smaller than a given threshold; when the bit error rate is underestimated, the effectiveness of channel coding drops dramatically and so does the decoded video quality. In this paper, we propose a low-complex, scalable video coding architecture based on compressive sensing (SVCCS) for wireless unicast and multicast transmissions. SVCCS achieves good scalability, error resilience and coding efficiency. SVCCS encoded bitstream is divided into base and enhancement layer. The layered structure provides quality and temporal scalability. While in the enhancement layer, the CS measurements provide fine granular quality scalability. In addition, we incorporate state-of-the-art technologies of compressive sensing to improve the coding efficiency. Experimental results show that SVCCS is more effective and efficient for wireless videocast than the existing solutions.
Siyuan Xiang, Lin Cai 0001
ICC1
2010 Distortion Analysis of Wyner-Ziv Distributed Video Coding
abstract
The Distributed Video Coding (DVC) follows an approach different from the conventional video coding. DVC has a simpler encoder but a more complicated decoder. This feature makes it possible to encode video in computation and energy constrained devices. Thus, DVC is appealing in sensor networks and other wireless networks. When transmitting real-time DVC encoded video streams, in order to adjust coding parameters according to the time-varying communication channel conditions or the dynamics of available bandwidth in the bottleneck, the source needs an efficient way to know the tradeoff of the coding parameters and the decoded video quality. However, how to quantify the DVC video quality using tractable models is an open issue. In this paper, we propose a distortion analysis model for DVC encoded Wyner-Ziv frames. The proposed closed-form distortion model for Wyner-Ziv frames is based on the reconstruction method of the "nearest neighbor binning". With the distortion analysis model, the average video frame PSNR can be estimated as a function of the codec parameters and the video statistics. Extensive simulations with different types of videos have been conducted and the results validate the accuracy of the proposed model. The model will be an enabling tool to further optimize the system parameters and network protocols for supporting DVC coded video over wireless and wired networks.
Siyuan Xiang, Lin Cai 0001
GLOBECOM1
2010 Scalable Modulation for Scalable Wireless Videocast
abstract
In conventional wireless systems with layered architectures, the physical layer treats all data streams from upper layers equally and apply the same modulation and coding schemes. Newer systems such as Digital Video Broadcast start to introduce hierarchical modulation schemes with SuperPosition preCoding (SPC) and support data streams of different priorities. However, SPC requires specialized hardware and has high complexity beyond most existing handheld devices. We thus propose scalable modulation (s-mod) by reusing the current mainstream modulation schemes with software-based bit-remapping. In this paper, we study how to optimize the configuration of the PHY layer s-mod and coding schemes to maximize the utility of videos with Scalable Video Coding (SVC). Simulation results demonstrate significant performance gains using s-mod and the cross-layer optimization, indicating s-mod and SVC is a good combination for wireless video multicast.
Lin Cai 0001, Yuanqian Luo, Siyuan Xiang, Jianping Pan 0001
INFOCOM3