notesonly.in

One notebook for every subject — open it anywhere.

Log in

Big data basics: Hadoop, Spark introduction

Data Science · Engineering

Study notes

Process 10TB logs: HDFS splits across 100 nodes; MapReduce maps (parse lines) then reduces (count errors) in parallel. Spark caches in memory: iterative ML 100x faster than MapReduce's disk writes. A 3-node Spark cluster counts billions of rows in minutes.

← Back to topics for Engineering