IIT Madras Big Data Lectures: Introduction to Big Data
BSDA500160 lectures10 weeks
Weeks
- Week 16 lecturesIntroduction to big data fundamentals · Big data in the era of cloud evolution · Big data – key concepts divide & conquer strategy · Big data – key concepts (contd) map & reduce functions · Spark introduction apache spark · Duality of sorting & hashing big data hashing
- Week 25 lecturesDifferent facets of cloud · Cloud native architecture principles · Services in public cloud · Types of cloud services · Types of cloud services (contd) & big data in era of cloud
- Week 36 lecturesData source types | data pipelines | big data engineering · Varieties of data formats (e.g. CSV) · Processing data formats relational vs non relational data · Software & data facts dimensions · Data lake definition vs data warehouse data mart · Extracting data from sources snapshotting
- Week 49 lecturesHadoop overview · Spark overview spark sql streaming · Architecture of spark cluster · Spark internals · Spark SQL · Hands-on creating a dataproc cluster · Hands-on working with real world data github · Hands-on spark sql data processing · Hands-on data processing using dataframe operations
- Week 55 lecturesExtract, transform, load (ETL) definition · ETL patterns keys processing reference data 1 · ETL patterns operational metadata data · Data lifecycle management pattern · Extract, transform, load (ELT) & scheduler
- Week 66 lecturesIntroduction to SQL relational algebra declarative programming pyspark · SQL planning & execution – spark SQL optimization, logical & physical plans · Types of joins – nested loops, merge join, hash join & big data variants · SQL optimization with explain & analyze – spark SQL & big data · Optimal execution in spark SQL – cost-based optimization · Introduction to NoSQL databases
- Week 76 lecturesMotivation for streaming – batch vs. streaming · Streaming context part 1 · Streaming applications & best practices – performance, state & fault tolerance · Event streaming & google pub/sub_1 · Google pub/sub : use cases, scalability & key properties · Google dataflow & apache beam – batch, streaming & pipelines
- Week 87 lecturesIntroduction to apache kafka – event streaming & architecture · Kafka design walkthrough – partitions, replication & offsets · Message semantics · Transactions for exactly once · Introduction to spark streaming | structured streaming & real-time data processing · Time-based event processing | windowing, triggers & streaming analytics · Hands-on with structured streaming | real-time data processing in jupyter notebook
- Week 96 lecturesBig data machine learning on a single machine | shared Memory & dask · Machine learning on a cluster · Introduction to spark mllib | scalable machine learning with pipelines & automl · Illustrative examples of spark mlLlib · Introduction to mlops · Best practices in ML on big data
- Week 104 lecturesIntroduction to deep learning on big data · Distributed deep learning with horovods · Deep learning scoring on streaming data · Big data course recap | key concepts, final exam prep & career insights