Introduction to the hadoop ecosystem by Uwe Seiler

Introduction to the Hadoop ecosystem

About me

About us

Why Hadoop?

How to scale data?

w1 w2 w3

r1 r2 r3

But…

What is Hadoop?

The Hadoop App Store

HDFS MapRed HCat Pig Hive HBase Ambari Avro Cassandra

Chukwa

Flume Hana HyperT Impala Mahout Nutch Oozie Scoop

Scribe Tez Vertica Whirr ZooKee Cloudera Horton MapR EMC

IBM Talend TeraData Pivotal Informat Microsoft. Pentaho Jasper

Kognitio Tableau Splunk Platfora Rack Karma Actuate MicStrat

Data Storage

Hadoop Distributed File System

HDFS Architecture

Data Processing

MapReduce

Typical large-data problem

MapReduce Flow

𝐤𝟏 𝐯𝟏 𝐤𝟐 𝐯𝟐 𝐤𝟒 𝐯𝟒 𝐤𝟓 𝐯𝟓 𝐤𝟔 𝐯𝟔 𝐤𝟑 𝐯𝟑

a 𝟏 b 2 c 9 a 3 c 2 b 7 c 8

a 𝟏 b 2 c 3 c 6 a 3 c 2 b 7 c 8

a 1 3 b 𝟐 7 c 2 8 9

a 4 b 9 c 19

Jobs & Tasks

Combined Hadoop Architecture

Word Count Mapper in Java

public class WordCountMapper extends MapReduceBase implements

Mapper<LongWritable, Text, Text, IntWritable>

private final static IntWritable one = new IntWritable(1);

private Text word = new Text();

public void map(LongWritable key, Text value, OutputCollector<Text,

IntWritable> output, Reporter reporter) throws IOException

String line = value.toString();

StringTokenizer tokenizer = new StringTokenizer(line);

while (tokenizer.hasMoreTokens())

word.set(tokenizer.nextToken());

output.collect(word, one);

Word Count Reducer in Java

public class WordCountReducer extends MapReduceBase

implements Reducer<Text, IntWritable, Text, IntWritable>

public void reduce(Text key, Iterator values, OutputCollector

output, Reporter reporter) throws IOException

int sum = 0;

while (values.hasNext())

IntWritable value = (IntWritable) values.next();

sum += value.get();

output.collect(key, new IntWritable(sum));

Scripting for Hadoop

Apache Pig

••

Pig in the Hadoop ecosystem

Distributed Programming Framework

Metadata Management

Scripting

Pig Latin

users = LOAD 'users.txt' USING PigStorage(',') AS (name,

pages = LOAD 'pages.txt' USING PigStorage(',') AS (user,

filteredUsers = FILTER users BY age >= 18 and age <=50;

joinResult = JOIN filteredUsers BY name, pages by user;

grouped = GROUP joinResult BY url;

summed = FOREACH grouped GENERATE group,

COUNT(joinResult) as clicks;

sorted = ORDER summed BY clicks desc;

top10 = LIMIT sorted 10;

STORE top10 INTO 'top10sites';

Pig Execution Plan

Try that with Java…

SQL for Hadoop

Apache Hive

Hive in the Hadoop ecosystem

Distributed Programming Framework

Metadata Management

Scripting Query

Hive Architecture

Hive Example

CREATE TABLE users(name STRING, age INT);

CREATE TABLE pages(user STRING, url STRING);

LOAD DATA INPATH '/user/sandbox/users.txt' INTO

TABLE 'users';

LOAD DATA INPATH '/user/sandbox/pages.txt' INTO

TABLE 'pages';

SELECT pages.url, count(*) AS clicks FROM users JOIN

pages ON (users.name = pages.user)

WHERE users.age >= 18 AND users.age <= 50

GROUP BY pages.url

SORT BY clicks DESC

LIMIT 10;

Bringing it all together…

Online Advertising

Getting started…

Hortonworks Sandbox

Hadoop Training

••

The end…or the beginning?

Introduction to the hadoop ecosystem by Uwe Seiler

Technology

SEILER John SL414 - E. C. Saylor · John "Hans" SEILER and Elizabeth BLOUGH 1. John "Hans"1 SEILER [SL414+], born i, 26 Aug 1768 in Pennsylvania; died i 4 Mar 1855, son of John SEILER

Trimble Business Center - Seiler-Geospatial

Uwe Seiler, Data Architect and Trainer at codecentric AG - "Hadoop & Germany & 2016"

Shana Lutker at Barbara Seiler Gallery

Seiler - Search Cost and BLP

Design: Mathias Seiler

Stepen Seiler - Training periodization in endurance sport

UWE 350 UWE 600 UWE 900 - KARL DEUTSCH

SEILER EPSTEIN ZIEGLER APPLEGATE LLP

Hadoop , Hadoop , Hadoop !!!

SEILER 202/402

Panelists Marc Seiler,University of Kaiserslautern, Germany

About Seiler Instrument Page 4

Session 62 Andreas Seiler

Uwe Habermann Uwe@VandU.eu Venelina Jordanova Venelina@VandU.eu Wishlist Silverswitch

FRANK J. SEILER RESEARCH LABORATORY

A commitment to excellence since 1945 - Seiler Inst...g Optical instruments have been a Seiler family tradition since 1913 when company founder Eric H. Seiler entered the ZEISS School

Seiler Ammonia and Alzheimer's Disease

UWE Bristol

SEILER IQ · SEILER IQ Dental Microscope LED. Table of Contents ... The P1026 product is an LED illumination system integrated into a Seiler surgical microscope. The power supply