Big data basics: Hadoop, Spark introduction
Data Science · Engineering
Study notes
Process 10TB logs: HDFS splits across 100 nodes; MapReduce maps (parse lines) then reduces (count errors) in parallel. Spark caches in memory: iterative ML 100x faster than MapReduce's disk writes. A 3-node Spark cluster counts billions of rows in minutes.