Taming Big Data with Apache Spark and Python

Taming Big Data with Apache Spark and Python

5 Hours
Deal Price$24.00
Suggested Price
$89.00
You save 73%
Taming Big Data with Apache Spark and Python
$24.00$89.0073% OFF
Taming Big Data with Apache Spark and Python

46 Lessons (5h)

  • Getting Started with Spark
    Introduction
    [Activity] Installing Enthought Canopy
    [Activity] Installing a JDK
    [Activity] Installing Spark
    [Activity] Installing the MovieLens Movie Rating Dataset3:35
    [Activity] Run your first Spark program! Ratings histogram example.
  • Spark Basics and Simple Examples
    Introduction to Spark
    The Resilient Distributed Dataset (RDD)
    Ratings Histogram Walkthrough
    Key/Value RDD's, and the Average Friends by Age Example
    [Activity] Running the Average Friends by Age Example
    Filtering RDD's, and the Minimum Temperature by Location Example
    [Activity]Running the Minimum Temperature Example, and Modifying it for Maximums
    [Activity] Running the Maximum Temperature by Location Example
    [Activity] Counting Word Occurrences using flatmap()
    [Activity] Improving the Word Count Script with Regular Expressions
    [Activity] Sorting the Word Count Results7:44
    [Exercise] Find the Total Amount Spent by Customer4:01
    [Excercise] Check your Results, and Now Sort them by Total Amount Spent.5:08
    Check Your Sorted Implementation and Results Against Mine.
  • Advanced Examples of Spark Programs
    [Activity] Find the Most Popular Movie
    [Activity] Use Broadcast Variables to Display Movie Names Instead of ID Numbers
    Find the Most Popular Superhero in a Social Graph
    [Activity] Run the Script - Discover Who the Most Popular Superhero is!
    Superhero Degrees of Separation: Introducing Breadth-First Search
    Superhero Degrees of Separation: Accumulators, and Implementing BFS in Spark
    [Activity] Superhero Degrees of Separation: Review the Code and Run it
    Item-Based Collaborative Filtering in Spark, cache(), and persist()
    [Activity] Running the Similar Movies Script using Spark's Cluster Manager
    [Exercise] Improve the Quality of Similar Movies
  • Running Spark on a Cluster
    Introducing Elastic MapReduce
    [Activity] Setting up your AWS / Elastic MapReduce Account and Setting Up PuTTY
    Partitioning
    Create Similar Movies from One Million Ratings - Part 1
    [Activity] Create Similar Movies from One Million Ratings - Part 2
    Create Similar Movies from One Million Ratings - Part 3
    Troubleshooting Spark on a Cluster
    More Troubleshooting, and Managing Dependencies
  • Other Spark Technologies and Libraries
    Introducing MLLib
    [Activity] Using MLLib to Produce Movie Recommendations
    Analyzing the ALS Recommendations Results
    Spark SQL
    Spark Streaming and GraphX
  • You Made It! Where to Go from Here.
    Learning More about Spark and Data Science
Taming Big Data with Apache Spark and Python
$24.00$89.0073% OFF
View similar items
DescriptionInstructorImportant DetailsRelated Products

Learn the Techniques Used by Major Companies to Manage Mass Data Sets

FK
Frank KaneFrank Kane spent 9 years at Amazon and IMDb, developing and managing the technology that automatically delivers product and movie recommendations to hundreds of millions of customers, all the time. Frank holds 17 issued patents in the fields of distributed computing, data mining, and machine learning. In 2012, Frank left to start his own successful company, Sundog Software, which focuses on virtual reality environment technology, and teaching others about big data analysis.

Description

Have you ever wondered how major companies and organizations manage all of the massive amounts of data they collect? The answer is Big Data technology, and Big Data engineers are in big-time demand. Major employers like Amazon, eBay, and NASA JPL use Apache Spark to extract data sets across a fault-tolerant Hadoop cluster. Sound complicated? That's why you should take this course, to learn these techniques and more, using your own system at home.

  • Access 46 lectures & 5 hours of content 24/7
  • Learn the concepts of Spark's Resilient Distributed Datastores
  • Develop & run Spark jobs quickly using Python
  • Translate complex analysis problems into iterative or multi-stage Spark scripts
  • Scale up to larger data sets using Amazon's Elastic MapReduce
  • Understand how Hadoop YARN distributes Spark across computing clusters
  • Learn about other Spark technologies, like Spark SQL, Spark Streaming, & GraphX

Specs

Details & Requirements

  • Length of time users can access this course: lifetime
  • Access options: web streaming
  • Certification of completion not included
  • Redemption deadline: redeem your code within 30 days of purchase
  • Experience level required: all levels

Terms

  • Unredeemed licenses can be returned for store credit within 30 days of purchase. Once your license is redeemed, all sales are final.
Your Cart
Your cart is empty. Continue Shopping!
Processing order...