Upload
vikran-kottaisamy
View
179
Download
1
Embed Size (px)
Citation preview
Data Preprocessing
BY:K.KOTTAISAMY
II MCA
Data cleaning
Data integration
Data transformation
Data reduction Data discretization
Data Preprocessing
Data cleaning Fill in missing values, smooth noisy data, identify or remove outliers, and resolve inconsistencies
Data integration Integration of multiple databases, data cubes, or files
Data transformation Normalization and aggregation
Major tasks in Data Preprocessing
Data reduction Reduced representation in volume but produces the same or similar analytical results
Data discretization Part of data reduction but with particular importance, especially for numerical data
Data Preprocessing
Data in the Real World Is Dirty: Lots of incorrect data, e.g., instrument faulty, human or computer error, transmission error
incomplete: lacking attribute values, containing only aggregate data
e.g., Occupation = “ ” (missing data) noisy: containing noise, errors, or outliers
e.g., Salary = “−10” (an error)
Data Cleaning
Data integration: combines data from multiple sources
Schema integration integrate metadata from different sources
Detecting and resolving data value conflicts for the same real world entity, attribute
values from different sources are different, e.g., different scales, metric vs. British units
Removing duplicates and redundant data
Data Integration
Smoothing: remove noise from data Aggregation: summarization, data cube
construction Generalization: concept hierarchy climbing Normalization: scaled to fall within a small,
specified range min-max normalization z-score normalization normalization by decimal scaling
Data Transformation
Min-max normalization: to [new_minA, new_maxA]
Data Transformation: Normalization
AAA
AA
A
minnewminnewmaxnewminmax
minvv _)__('
Ex. Let income range $12,000 to $98,000 normalized to [0.0, 1.0]. Then $73,000 is mapped to
716.00)00.1(000,12000,98
000,12600,73
Z-score normalization (μ: mean, σ: standard deviation):
Ex. Let μ = 54,000, σ = 16,000. Then
Data Transformation: Normalization
A
Avv
'
225.1000,16
000,54600,73
Normalization by decimal scaling
Where j is the smallest integer such that Max(|ν’|) < 1
Data Transformation: Normalization
j
vv
10'
Why data reduction? A database/data warehouse may store terabytes of
data
Complex data analysis/mining may take a very long time to run on the complete data set
Data reduction The data set that is much smaller in volume but yet
produce the same analytical results
Data reduction
Data reduction strategies
Data cube aggregation
Dimensionality reduction — e.g., remove unimportant attributes
Data Compression
Discretization and concept hierarchy generation
The lowest level of a data cube (base cuboid)
The aggregated data for an individual entity of interest
E.g., a customer in a phone calling data warehouse
Multiple levels of aggregation in data cubes
Further reduce the size of data to deal with
Data cube aggregation
Feature selection (attribute subset selection): Select a minimum set of attributes that is
sufficient for the data mining task. Heuristic methods
step-wise forward selection step-wise backward elimination combining forward selection and backward
elimination
Dimensionality Reduction
Three types of attributes: Nominal — values from an unordered set Ordinal — values from an ordered set Continuous — real numbers
Discretization: Some classification algorithms only accept
categorical attributes. Reduce data size by discretization Prepare for further analysis
Data Discretization
THANKING YOU