Skip to content
September 2026 qualifier: applications close Sun 27 Sep · Week 1 starts Fri 2 Oct
Qualifier Hub

IIT Madras Big Data Lectures: Introduction to Big Data

BSDA500160 lectures10 weeks

Weeks

  1. Week 16 lecturesIntroduction to big data fundamentals · Big data in the era of cloud evolution · Big data – key concepts divide & conquer strategy · Big data – key concepts (contd) map & reduce functions · Spark introduction apache spark · Duality of sorting & hashing big data hashing
  2. Week 25 lecturesDifferent facets of cloud · Cloud native architecture principles · Services in public cloud · Types of cloud services · Types of cloud services (contd) & big data in era of cloud
  3. Week 36 lecturesData source types | data pipelines | big data engineering · Varieties of data formats (e.g. CSV) · Processing data formats relational vs non relational data · Software & data facts dimensions · Data lake definition vs data warehouse data mart · Extracting data from sources snapshotting
  4. Week 49 lecturesHadoop overview · Spark overview spark sql streaming · Architecture of spark cluster · Spark internals · Spark SQL · Hands-on creating a dataproc cluster · Hands-on working with real world data github · Hands-on spark sql data processing · Hands-on data processing using dataframe operations
  5. Week 55 lecturesExtract, transform, load (ETL) definition · ETL patterns keys processing reference data 1 · ETL patterns operational metadata data · Data lifecycle management pattern · Extract, transform, load (ELT) & scheduler
  6. Week 66 lecturesIntroduction to SQL relational algebra declarative programming pyspark · SQL planning & execution – spark SQL optimization, logical & physical plans · Types of joins – nested loops, merge join, hash join & big data variants · SQL optimization with explain & analyze – spark SQL & big data · Optimal execution in spark SQL – cost-based optimization · Introduction to NoSQL databases
  7. Week 76 lecturesMotivation for streaming – batch vs. streaming · Streaming context part 1 · Streaming applications & best practices – performance, state & fault tolerance · Event streaming & google pub/sub_1 · Google pub/sub : use cases, scalability & key properties · Google dataflow & apache beam – batch, streaming & pipelines
  8. Week 87 lecturesIntroduction to apache kafka – event streaming & architecture · Kafka design walkthrough – partitions, replication & offsets · Message semantics · Transactions for exactly once · Introduction to spark streaming | structured streaming & real-time data processing · Time-based event processing | windowing, triggers & streaming analytics · Hands-on with structured streaming | real-time data processing in jupyter notebook
  9. Week 96 lecturesBig data machine learning on a single machine | shared Memory & dask · Machine learning on a cluster · Introduction to spark mllib | scalable machine learning with pipelines & automl · Illustrative examples of spark mlLlib · Introduction to mlops · Best practices in ML on big data
  10. Week 104 lecturesIntroduction to deep learning on big data · Distributed deep learning with horovods · Deep learning scoring on streaming data · Big data course recap | key concepts, final exam prep & career insights

Big Data notesBig Data previous year papers