Course Hive
Search

Welcome

Sign in or create your account

Continue with Google
or
30 Data Skipping and Z-Ordering in Delta Lake Tables | Optimize & Data Compaction Delta Lake Tables
Play lesson

PySpark - Zero to Hero | PySpark Tutorial 2025 | Spark Tutorial 2025 | Learn from Basics to Advanced Performance Optimization - 30 Data Skipping and Z-Ordering in Delta Lake Tables | Optimize & Data Compaction Delta Lake Tables

4.0 (1)
18 learners

What you'll learn

This course includes

  • 9 hours of video
  • Certificate of completion
  • Access on mobile and TV

Summary

Keywords

Full Transcript

Video explains - What is the impact of data skipping on jobs? How z-ordering in delta lake works ? How to optimize delta lake tables? Chapters 00:00 - Introduction 00:31 - What is Data Skipping and Z-Ordering in Delta Lake? 03:34 - Z-Ordering for more than 1 column/Multidimensional Z-ORDER 04:38 - Delta Lake Table Optimization with Example 11:59 - Multi Column Z-Ordering in Delta Lake Table 14:43 - Impact of Partitioning with Z-Ordering 16:24 - Selective Z-Ordering with Partition filters 17:57 - Auto Compaction in Delta Lake Table For Local PySpark Jupyter Lab setup just run the command - docker pull jupyter/pyspark-notebook Python Basics - https://www.learnpython.org/ GitHub URL for code - https://github.com/subhamkharwal/pyspark-zero-to-hero/blob/master/25_delta_lake_optimization_and_z_ordering.ipynb Delta Lake Optimization Documentation - https://docs.delta.io/latest/optimizations-oss.html#language-sql The series provides a step-by-step guide to learning PySpark, a popular open-source distributed computing framework that is used for big data processing. New video in every 3 days ❤️ #spark #pyspark #python #dataengineering

Course Hive

Continue this lesson in the app

Install CourseHive on Android or iOS to keep learning while you move.

Related Courses

FAQs

Course Hive
Download CourseHive
Keep learning anywhere