Running Presto with Alluxio on Amazon EMR
February 12, 2020
By 
Alex Ma

Many organizations are leveraging EMR to run big data analytics on public cloud. However, reading and writing data to S3 directly can result in slow and inconsistent performance. Alluxio is a data orchestration layer for the cloud, and in this use case it caches data for S3, ensuring high and predictable performance as well as reduced network traffic.

In this office hour, you will learn about:

  • How to set up Alluxio with the EMR stack so that Presto jobs can seamlessly read from and write to S3
  • Compare the performance between Presto on EMR with Presto and Alluxio on EMR
  • Open Session for discussion on any topics such as solving the separation of compute and storage problem, and more
ALLUXIO COMMUNITY OFFICE HOUR

Many organizations are leveraging EMR to run big data analytics on public cloud. However, reading and writing data to S3 directly can result in slow and inconsistent performance. Alluxio is a data orchestration layer for the cloud, and in this use case it caches data for S3, ensuring high and predictable performance as well as reduced network traffic.

In this office hour, you will learn about:

  • How to set up Alluxio with the EMR stack so that Presto jobs can seamlessly read from and write to S3
  • Compare the performance between Presto on EMR with Presto and Alluxio on EMR
  • Open Session for discussion on any topics such as solving the separation of compute and storage problem, and more

Video:

Slides:

Running Presto with Alluxio on Amazon EMR from Alluxio, Inc.

Complete the form below to access the full overview:

Videos

Sign-up for a Live Demo or Book a Meeting with a Solutions Engineer