Taming Big Data with Apache Spark and Python
5 Hours
Deal Price$24.00
Suggested Price
$89.00
You save 73%
Taming Big Data with Apache Spark and Python
$24.00$89.0073% OFF
46 Lessons (5h)
- Getting Started with SparkIntroduction[Activity] Installing Enthought Canopy[Activity] Installing a JDK[Activity] Installing Spark[Activity] Installing the MovieLens Movie Rating Dataset3:35[Activity] Run your first Spark program! Ratings histogram example.
- Spark Basics and Simple ExamplesIntroduction to SparkThe Resilient Distributed Dataset (RDD)Ratings Histogram WalkthroughKey/Value RDD's, and the Average Friends by Age Example[Activity] Running the Average Friends by Age ExampleFiltering RDD's, and the Minimum Temperature by Location Example[Activity]Running the Minimum Temperature Example, and Modifying it for Maximums[Activity] Running the Maximum Temperature by Location Example[Activity] Counting Word Occurrences using flatmap()[Activity] Improving the Word Count Script with Regular Expressions[Activity] Sorting the Word Count Results7:44[Exercise] Find the Total Amount Spent by Customer4:01[Excercise] Check your Results, and Now Sort them by Total Amount Spent.5:08Check Your Sorted Implementation and Results Against Mine.
- Advanced Examples of Spark Programs[Activity] Find the Most Popular Movie[Activity] Use Broadcast Variables to Display Movie Names Instead of ID NumbersFind the Most Popular Superhero in a Social Graph[Activity] Run the Script - Discover Who the Most Popular Superhero is!Superhero Degrees of Separation: Introducing Breadth-First SearchSuperhero Degrees of Separation: Accumulators, and Implementing BFS in Spark[Activity] Superhero Degrees of Separation: Review the Code and Run itItem-Based Collaborative Filtering in Spark, cache(), and persist()[Activity] Running the Similar Movies Script using Spark's Cluster Manager[Exercise] Improve the Quality of Similar Movies
- Running Spark on a ClusterIntroducing Elastic MapReduce[Activity] Setting up your AWS / Elastic MapReduce Account and Setting Up PuTTYPartitioningCreate Similar Movies from One Million Ratings - Part 1[Activity] Create Similar Movies from One Million Ratings - Part 2Create Similar Movies from One Million Ratings - Part 3Troubleshooting Spark on a ClusterMore Troubleshooting, and Managing Dependencies
- Other Spark Technologies and LibrariesIntroducing MLLib[Activity] Using MLLib to Produce Movie RecommendationsAnalyzing the ALS Recommendations ResultsSpark SQLSpark Streaming and GraphX
- You Made It! Where to Go from Here.Learning More about Spark and Data Science
Taming Big Data with Apache Spark and Python
$24.00$89.0073% OFF
DescriptionInstructorImportant DetailsRelated Products
Learn the Techniques Used by Major Companies to Manage Mass Data Sets
FK
Frank KaneFrank Kane spent 9 years at Amazon and IMDb, developing and managing the technology that automatically delivers product and movie recommendations to hundreds of millions of customers, all the time. Frank holds 17 issued patents in the fields of distributed computing, data mining, and machine learning. In 2012, Frank left to start his own successful company, Sundog Software, which focuses on virtual reality environment technology, and teaching others about big data analysis.Terms
- Unredeemed licenses can be returned for store credit within 30 days of purchase. Once your license is redeemed, all sales are final.
Your Cart
Your cart is empty. Continue Shopping!
Processing order...


