Presto: Fast SQL-on-anything across data lakes, DBMS, and NoSQL Data stores

December 13, 2020

Kamil Bajda-Pawlikowski

CTO

Starburst

Presto, an open source distributed SQL engine, is widely recognized for its low-latency queries, high concurrency, and native ability to query multiple data sources. Proven at scale in a variety of use cases at Comcast, GrubHub, FINRA, LinkedIn, Lyft, Netflix, Slack, Zalando, in the last few years Presto experienced an unprecedented growth in popularity in both on-premises and cloud deployments over Object Stores, HDFS, NoSQL and RDBMS data stores.

Delta Lake, a storage layer originally invented by Databricks and recently open sourced, brings ACID capabilities to big datasets held in Object Storage. While initially designed for Spark, Delta Lake now supports multiple query compute engines including Presto.

In this talk we discuss how Presto enables query-time correlations between Delta Lake, Snowflake, and Elasticsearch to drive interactive BI analytics across disparate datasets.

Video:

Presentation Slides:

Presto: Fast SQL-on-Anything Across Data Lakes, DBMS, and NoSQL Data Stores from Alluxio, Inc.

‍

In this talk we discuss how Presto enables query-time correlations between Delta Lake, Snowflake, and Elasticsearch to drive interactive BI analytics across disparate datasets.

Videos:

Presentation Slides:

Presto: Fast SQL-on-anything across data lakes, DBMS, and NoSQL Data stores from Alluxio, Inc.

Complete the form below to access the full overview:

Videos

Tech Talk: How Coupang Leverages Distributed Cache to Accelerate ML Model Training

Coupang is a leading e-commerce company in South Korea, with over 50,000 employees and $20+ billion in annual revenue. Coupang's AI platform team builds and manages a large-scale AI platform in AWS for machine learning engineers to train models that enhance and customize product search results and product recommendations for its 100+ million customers.

As the search and recommendation models evolve, optimizing the underlying infrastructure for AI/ML workloads is essential for the e-commerce business. Coupang's platform team actively sought to improve their model training pipeline to boost machine learning engineers' productivity, publish models to production faster, and reduce operational costs.

Coupang focused on addressing several key areas:

Shortening data preparation and model training time
Improving GPU utilization in training clusters in different regions
Reducing S3 API and egress costs incurred from copying large training datasets across regions
Simplifying the operational complexity of storage system management

In this tech talk, Hyun Jung Baek, Staff Backend Engineer at Coupang, will share best practices for leveraging distributed caching to power search and recommendation model training infrastructure.

Hyun will discuss:

How Coupang builds a world-class large-scale AI platform for machine learning engineers to deliver better search and recommendation models
How adding distributed caching to their multi-region AI infrastructure improves GPU utilization, accelerates end-to-end training time, and significantly reduces cross-region data transfer costs.
How to simplify platform operations and to easily deploy the same architecture to new GPU clusters.

About the Speaker

Hyun Jung Baek is a Staff Backend Engineer at Coupang.

‍

April 22, 2025

GTC 2025 | Alluxio Decouples Storage and Compute for a Faster AI Future

April 9, 2025

Inside Deepseek 3FS: A Deep Dive into AI-Optimized Distributed Storage

Deepseek’s recent announcement of the Fire-flyer File System (3FS) has sparked excitement across the AI infra community, promising a breakthrough in how machine learning models access and process data.

In this webinar, an expert in distributed systems and AI infrastructure will take you inside Deepseek 3FS, the purpose-built file system for handling large files and high-bandwidth workloads. We’ll break down how 3FS optimizes data access and speeds up AI workloads as well as the design tradeoffs made to maximize throughput for AI workloads.

This webinar you’ll learn about how 3FS works under the hood, including:

✅ The system architecture

✅ Core software components

✅ Read/write flows

✅ Data distribution/placement algorithms

✅ Cluster/node management and disaster recovery

Whether you’re an AI researcher, ML engineer, or infrastructure architect, this deep dive will give you the technical insights you need to determine if 3FS is the right solution for you.

‍

April 1, 2025

Sign-up for a Live Demo or Book a Meeting with a Solutions Engineer

Request a demo