Course Hive
Search

Welcome

Sign in or create your account

Continue with Google
or
29 Optimize Data Scanning with Partitioning in Spark | How Partitioning data works | Optimize Jobs
Play lesson

PySpark - Zero to Hero | PySpark Tutorial 2025 | Spark Tutorial 2025 | Learn from Basics to Advanced Performance Optimization - 29 Optimize Data Scanning with Partitioning in Spark | How Partitioning data works | Optimize Jobs

4.0 (1)
18 learners

What you'll learn

This course includes

  • 9 hours of video
  • Certificate of completion
  • Access on mobile and TV

Summary

Keywords

Full Transcript

Video explains - What is the impact of data scanning on jobs? How partitioning works ? How to avoid un-necessary data scanning? How to optimize jobs using Partitioning? Chapters 00:00 - Introduction 00:35 - Why avoiding un-necessary Data Scanning in Important? 02:12 - Impacat of Data Partitioning 03:07 - Imapct of partitioning column missing from Query 03:33 - Impact of parititioning on High Cardinality column 04:42 - Live data Examples For Local PySpark Jupyter Lab setup just run the command - docker pull jupyter/pyspark-notebook Python Basics - https://www.learnpython.org/ GitHub URL for code - https://github.com/subhamkharwal/pyspark-zero-to-hero/blob/master/24_data_scanning_and_partitioning.ipynb The series provides a step-by-step guide to learning PySpark, a popular open-source distributed computing framework that is used for big data processing. New video in every 3 days ❤️ #spark #pyspark #python #dataengineering

Course Hive

Continue this lesson in the app

Install CourseHive on Android or iOS to keep learning while you move.

Related Courses

FAQs

Course Hive
Download CourseHive
Keep learning anywhere