Hassan Shojania

dblp:53/4362 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-authorComputer networks · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 28% GPUs and heterogeneous computing · 24% Storage systems · 24%
Theoretical computer science
1 paper
Coding theory · 100%
Computer graphics and multimedia
1 paper
Multimedia systems and quality of experience · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Coding theory
network coding
0.112010
Tenor: making coding practical from servers to smartphones · ACM Multimedia 2010
GPUs and heterogeneous computing
GPU computing
0.112009
Nuclei: GPU-Accelerated Many-Core Network Coding · INFOCOM 2009
Parallel and multicore computing › many-core systems
many-core computing
0.112009
Nuclei: GPU-Accelerated Many-Core Network Coding · INFOCOM 2009
Storage systems › storage reliability › erasure coding
network coding
0.112009
Nuclei: GPU-Accelerated Many-Core Network Coding · INFOCOM 2009
Multimedia systems and quality of experience
multimedia streaming
0.012001
Experiences with MPEG-4 multimedia streaming · ACM Multimedia 2001
Multimedia systems and quality of experience
quality of service
0.012001
Experiences with MPEG-4 multimedia streaming · ACM Multimedia 2001

Methods — techniques the papers use, named apart from their topics

reed-solomon codes · 0.2random network coding · 0.2fountain codes · 0.2multi-threaded implementation · 0.1SIMD vector instructions · 0.1bandwidth measurement · 0.0UDP streaming · 0.0
YearPublicationVenuePosition
2010 Tenor: making coding practical from servers to smartphones
abstract
It has been theoretically shown that performing coding in networked systems, including Reed-Solomon codes, fountain codes, and random network coding, has a clear advantage with respect to simplifying the design of protocols. These coding techniques can be deployed on a wide range of networked nodes, from servers in the "cloud" to smartphone devices. However, large-scale real-world deployment of systems using coding is still rare, mainly due to the computational complexity of coding algorithms. This is especially a concern on both extremes: in high-bandwidth servers where coding may not be able to saturate the uplink bandwidth, and in smartphone devices where hardware limitations prevail.
Hassan Shojania, Baochun Li
ACM Multimedia1
2009 Pushing the Envelope: Extreme Network Coding on the GPU
abstract
While it is well known that network coding achieves optimal flow rates in multicast sessions, its potential for practical use has remained to be a question, due to its high computational complexity. With GPU computing gaining momentum as a result of increased hardware capabilities and improved programmability, we show in this paper how the GPU can be used to improve network coding performance dramatically. Our previous work presented the first attempt in the literature to maximize the performance of network coding by taking advantage of not only multi-core CPUs, but also hundreds of computing cores in commodity off-the-shelf Graphics Processing Units (GPU). This paper represents another step forward, and presents a new array of GPU-based algorithms that improve network encoding by a factor of 2.2, and network decoding by a factor of 2.7 to 27.6 across a range of practical configurations. With just a single NVIDIA GTX 280 GPU, our implementation of GPU-based network encoding outperforms an 8-core Intel Xeon server by a margin of at least 4.3 to 1 in all practical test cases, and over 3000 peers can be served at high-quality video rates if network coding is used in a streaming server. With 128 blocks, for example, coding rates up to 294 MB/second can be achieved with a variety of block sizes.
Hassan Shojania, Baochun Li
ICDCS1
2009 Nuclei: GPU-Accelerated Many-Core Network Coding
abstract
While it is a well known result that network coding achieves optimal flow rates in multicast sessions, its potential for practical use has remained to be a question, due to its high computational complexity. Our previous work has attempted to design a hardware-accelerated and multi-threaded implementation of network coding to fully utilize multi-core CPUs, as well as SSE2 and AltiVec SIMD vector instructions on X86 and PowerPC processors. This paper represents another step forward, and presents the first attempt in the literature to maximize the performance of network coding by taking advantage of not only multi-core CPUs, but also potentially hundreds of computing cores in commodity off-the-shelf graphics processing units (GPU). With GPU computing gaining momentum as a result of increased hardware capabilities and improved programmability, our work shows how the GPU, with a design involving thousands of lightweight threads, can boost network coding performance significantly. Many-core GPUs can be deployed as an attractive alternative and complementary solution to multi-core servers, by offering a better price/performance advantage. In fact, multi-core CPUs and many-core GPUs can be deployed and used to perform network coding simultaneously, potentially useful in media streaming servers where hundreds of peers are served concurrently by these dedicated servers. In this paper, we present Nuclei, the design and implementation of GPU-based network coding. With Nuclei, only one mainstream NVidia 8800 GT GPU outperforms an 8-core Intel Xeon server in most test cases. A combined CPU-GPU encoding scenario achieves coding rates of up to 116 MB/second for a variety of coding settings, which is sufficient to saturate a Gigabit Ethernet interface.
Hassan Shojania, Baochun Li, Xin Wang 0002
INFOCOM1
2009 Random network coding on the iPhone: fact or fiction?
abstract
In multi-hop wireless networks, random network coding represents the general design principle of transmitting random linear combinations of blocks in the same "batch" to downstream relays or receivers. It has been recognized that random network coding in multi-hop wireless networks may improve unicast throughput in scenarios when multiple paths are simultaneously utilized between the source and the destination. However, the computational complexity of random network coding, and its energy consumption implications, may potentially limit its applicability and practicality in mobile devices. In this paper, we present our real-world implementation of random network coding on the Apple iPhone and iPod Touch mobile platforms, and offer an in-depth investigation with respect to the difficulties towards such an implementation, the limitations of the ARM processor and the hardware platform, as well as our hand-tuning efforts to maximize coding performance on the iPhone platform. With our implementation deployed on both the iPhone 3G and the second-generation iPod Touch, we report its coding performance, energy consumption rates, as well as CPU usage with multimedia streaming.
Hassan Shojania, Baochun Li
NOSSDAV1
2008 Crystal: An Emulation Framework for Practical Peer-to-Peer Multimedia Streaming Systems
abstract
To rapidly evolve new designs of peer-to-peer (P2P) multimedia streaming systems, it is highly desirable to test and troubleshoot them in a controlled and repeatable experimental environment in a local cluster of servers, as it is risky to integrate untested protocols in live production and mission-critical peer-to-peer sessions, such as live P2P streaming. Though it is possible to construct such controlled experiments with virtual machine monitors, there are a number of challenges and roadblocks: (1) The deployment of such resource-hungry virtual machine environments are complicated and time-consuming for researchers without prior systems expertise; (2) The system designer needs to implement many basic streaming elements, such as playback buffers and message switches. In this paper, we seek to address these challenges by introducing Crystal, an emulation framework for practical P2P multimedia streaming systems, which provides support for developing, testing, and troubleshooting new streaming system designs in a controlled server cluster environment. It is our imperative design objective that Crystal offers ease of use, rapid experimental turnaround, and the capability of emulating realistic P2P environments.
Mea Wang, Hassan Shojania, Baochun Li
ICDCS2
2007 Parallelized Progressive Network Coding With Hardware Acceleration
abstract
The fundamental insight of network coding is that information to be transmitted from the source in a session can be inferred, or decoded, by the intended receivers, and does not have to be transmitted verbatim. It is a well known result that network coding may achieve better network throughput in certain multicast topologies; however, the practicality of network coding has been questioned, due to its high computational complexity. This paper represents the first attempt towards a high performance implementation of network coding. We first propose to implement progressive decoding with Gauss-Jordan elimination, such that blocks can be decoded as they are received. We then employ hardware acceleration with SSE2 and AltiVec SIMD vector instructions on x86 and PowerPC processors, respectively. We then use a careful threading design to take advantage of symmetric multiprocessor (SMP) systems and multi-core processors. The objective of this work is to explore the computational limits of network coding in off-the-shelf modern processors, and to provide a solid reference implementation to facilitate commercial deployment of network coding. Our high-performance implementation is packaged as a C++ class library, and runs in Linux, Mac OS X and Windows, in Intel, AMD and IBM PowerPC processor families. On a Dual dual-core PowerPC G5 2.5 GHz server, the coding bandwidth of our implementation is able to reach 43 MB/second with 64 blocks of 32 KB each, achieving speedup of 21 over the baseline implementation.
Hassan Shojania, Baochun Li
IWQoS1
2006 Performance improvement of the H.264/AVC deblocking filter using SIMD instructions
abstract
The H.264/AVC standard defines an in-loop deblocking filter which is used in both the encoder and decoder. This work examines several methods for improving the performance of the H.264/AVC reference software implementation of the deblocking filter. Methods examined include general software optimization, parallelization through standard multimedia SIMD instructions, and augmenting standard SIMD instruction sets with new instructions. Using the above methods, we are able to achieve a large speedup of the deblocking filter computation.
Stephen Warrington, Hassan Shojania, Subramania Sudharsanan, Wai-Yip Chan
ISCAS2
2001 Experiences with MPEG-4 multimedia streaming
abstract
With the advent of next-generation multimedia technologies such as very-low bit rate MPEG-4 codec, multimedia streaming of high-quality video and audio has become a near-term reality. The high compression ratio and error resilience offered by the MPEG-4 standard promise near-term popularity for rich contents and exceptional quality to consumers over affordable Internet connections, such as xDSL, cable modem and 3G wireless networks. Audio and video streaming applications are at the center of such scenarios; and Quality-of-Service (QoS) support in such applications is critical to their widespread acceptance.To the best of our knowledge, there has been no existing open-source MPEG-4 multimedia streaming applications in the academic community, which leads to the lack of research results using MPEG-4 streaming, especially with respect to Quality-of-Service support. In this work, we have implemented an open-source MPEG-4 multimedia streaming testbed in IP-based networks. In this paper, we show our experiences and lessons learned with such a testbed. First, we describe the algorithms and solutions used in our implemention testbed, emphasizing several critical issues. Second, through extensive experiments, we demonstrate measurements of bandwidth requirements and data loss for streaming a set of multimedia samples with different bit rates over UDP, which is ubiquitously available in the TCP/IP protocol stack on all consumer operating systems. Finally, future work for further improvements is also discussed.
Hassan Shojania, Baochun Li
ACM Multimedia1