Exploratory Visual Analysis

Data Visualization

Prasanna Parasurama

Lifecycle of a data analysis project

Data acquisition

Provenance

  • Is the source reliable?
  • How was the data collected?
  • When was it collected?
  • What was the intent or motive?

Data understanding

  • What does the dataset contain?
  • What tables are available?
  • What does each row represent?
  • What data is missing?
  • Are there erroneous values?
  • Are the data types correct?

Data cleaning

  • What should we do with missing data?
  • Which data types need conversion?
  • What should we do with erroneous values?

Exploratory analysis

  • Explore variable distributions
  • Explore relationships in the data
  1. Ask a question / generate a hypothesis
  2. Create charts to answer the question
  3. Ask a new question based on the chart ↻ Repeat

Formal statistical modeling

Findings

Communication and dissemination

Our case study: Applicant Tracking System (ATS) data

We will work with Applicant Tracking System (ATS) data from a real company.

  • The ATS records all applications received by the company, information about each application, and the outcome of each application.
  • The level of observation is the application: each row corresponds to a unique application.
  • An applicant can submit multiple applications.
Applications Unique applicants Jobs Application dates
25,450 16,999 403 Jan 2016–Nov 2018

ATS data dictionary

Field Definition / values
application_id Unique application identifier.
applicant_id Unique applicant identifier; an applicant can have multiple applications.
app_date Date of the application.
job_id Unique job identifier.
biz_unit Business unit of the job.
gender Applicant gender: male, female.
race Applicant race: white, black, asian, hispanic.
yrs_exp Years of experience at the time of application.
highest_degree bachelors, masters, doctorate, diploma, n/a.
Field Definition / values
field_of_study technical, business, law, other.
school_rank top10, 11–20, 21–50, 51–100, 101–200, unranked.
school_region Region of the applicant’s school.
referral Referral from a current employee: 0 = no, 1 = yes.
job_resume_fit Resume skills’ fit to the job description: 0 = no fit, 1 = perfect fit.
callback Received a callback for this application: 0 = no, 1 = yes.
interview Received an interview for this application: 0 = no, 1 = yes.
offer Received an offer for this application: 0 = no, 1 = yes.
hired Hired for this application: 0 = no, 1 = yes.

What do we know about each application?

Field group Examples Questions we can explore
Application and job Application date, job, business unit Where is application volume concentrated?
Applicant background Experience, degree, field of study, school, gender, race How does the applicant mix vary?
Application signals Referral, job–resume fit (0–1) How do these relate to progression?
Outcomes Callback → interview → offer → hired Where do applications leave the process?

EDA for one variable

Describe the center, spread, and shape of a distribution.

Common EDA questions for 1 variable

Start by asking how the values of one field are distributed.

Question Summary Visual
What is the central tendency? Mean, median, mode Bar chart, histogram
How are values spread? Range, IQR, standard deviation Box plot, histogram
What is the shape? Skew, modes, tails Histogram, density, ECDF

Histograms visualize a distribution

Each bar counts applications in a one-year experience bin. Taller bars mean that experience level appears more often.

  • Read years of experience on the x-axis and application counts on the y-axis. Look for peaks, spread, and tails.

A density curve smooths the histogram

A density curve smooths the experience distribution. Peaks show concentrations; area over an interval approximates its share.

  • Height is density, not a count. The total area under the full curve is 1; smoothing can hide narrow peaks.

Mean: an estimate of central tendency

\(\bar{x}=\frac{1}{n}\sum_{i=1}^{n}x_i\). Add all experience values and divide by the number of applications.

An extreme value moves the mean

Most applications list substantially less experience than the maximum of 50 years.

Mean 10.62 (orange) · Median 9 (dashed teal)

Median: the middle ordered value

Half of applications have at most 9 years of experience; half have at least 9.

  • Sort the observations from smallest to largest.
  • Take the middle value, or average the two middle values when the count is even.
  • The median is also the 50th percentile.

Mode: the most frequent value

  • A distribution can have one mode, several modes, or no modes at all.
  • A category mode describes the most common category, not a numerical center.

Which years of experience are most common?

The tallest one-year bar is at 5 years, the mode of the experience distribution.

Percentiles: positions in an ordered distribution

A percentile is a cut point in an ordered distribution. The 75th percentile has about 75% of observations at or below it.

Read from smallest to largest:

123456789101112131415161718192015 of 20 values = 75% at or below 15

In ATS data, the 75th percentile is 14 years of experience.

Quartiles on the experience distribution

Quartiles divide ordered observations into four roughly equal parts: Q1 = 25th, Q2 = 50th (median), Q3 = 75th percentile.

Interquartile range

The IQR is the distance from the 25th to the 75th percentile.

\[\operatorname{IQR}=Q_3-Q_1=14-6=8\text{ years}\]

The middle half of ATS applications spans 6 to 14 years of experience.

Box plots summarize a distribution

The box marks Q1, median, and Q3; whiskers show a defined range, with extreme values shown separately.

Trimmed (or truncated) means

  • In diving, the highest and lowest scores are dropped so that individual judges cannot significantly influence the score.
  • Trimmed means can be thought of as a compromise between arithmetic means and medians.

Diving scores with two low and two high values crossed out, leaving 7, 7.5, and 7.5.

Outliers are not automatically errors

The highest recorded experience is 50 years.

  • It may describe a genuine applicant, a miscoded value, or a different definition of experience.
  • Removing it solely because it is unusual would hide a data-quality question.
  • Compare conclusions with and without extreme values, then document the decision.

Estimates of spread

Center does not tell us whether applications have similar or very different experience levels.

Measure Experience in ATS What it captures
Range 49 years Maximum − minimum
IQR 8 years Width of middle half
Standard deviation 6.67 years Typical squared-deviation scale

Range is the simplest spread estimate

Experience spans from 1 to 50 years across applications.

\[\operatorname{range}=\max(x)-\min(x)=50-1=49\text{ years}\]

The range depends on just two observations; it is sensitive to extremes.

Mean absolute deviation

Average how far observations lie from the mean, ignoring direction.

\[\operatorname{MAD}_{\bar{x}}=\frac{1}{n}\sum_{i=1}^{n}|x_i-\bar{x}|\]

For application experience, the mean absolute deviation is computed in years, the same unit as the original values.

Variance and standard deviation

Square deviations from the mean, average them, then take the square root to return to the original unit.

\[\operatorname{Var}(x)=\frac{1}{n}\sum_{i=1}^{n}(x_i-\bar{x})^2,\qquad \operatorname{SD}(x)=\sqrt{\operatorname{Var}(x)}\]

  • Variance has units of years².
  • Standard deviation has units of years.
  • Here we describe the observed applications, so the denominator is \(n\).

Standard deviation on the experience distribution

Experience has mean 10.62 years and standard deviation 6.67 years. The outer lines mark one SD below and above the mean.

Estimates of shape

After center and spread, inspect the form of the distribution.

  • Is it roughly symmetric or skewed?
  • Does it have one peak or several?
  • Are there long tails, gaps, or boundaries?

Skewness: the two tails need not match

The experience distribution has a longer right tail: a few applications report much more experience than most.

CDFs are useful to compare distributions

Cumulative distribution of individual income in 2013 for the world, China, and United States, with median incomes and an income threshold marked.

A cumulative distribution function (CDF) gives the share of observations at or below a value.

Read up from an income

What share has that income or less?

Read across from 50%

What is the median income?

The x-axis uses a logarithmic scale.

ECDFs compare fit scores across genders

At fit score x, each curve gives the share of applications with a fit score of x or less within that recorded gender.

EDA with two or more variables

Examine relationships and compare groups.

Choose a chart for the pair of variables

Pair of variables Question in ATS data Example chart
Numeric + numeric Does fit vary with experience? Small scatterplot of fit versus years of experience.
Numeric + category How does fit differ by unit? Small box plots of fit in two business units.
Category + category How does field-of-study mix vary by business unit? Heatmap comparing field-of-study shares across two business units.

Scatter plots for numeric-numeric comparisons

Fit tends to rise with experience, but applications with the same experience have very different fit scores.

Look for clusters and boundaries

Grouping by business unit can reveal different clouds.

Small multiples of scatterplots

Separate panels make the experience–fit relationship easier to inspect within each unit.

Bubble size adds a third quantity

Each point is a business unit; its area indicates the number of valid-fit applications.

Pearson correlation summarizes a linear relationship

Experience and fit have Pearson correlation 0.227.

\[-1\leq r\leq 1\]

  • Positive values describe an upward linear association.
  • Values near 0 describe little linear association.
  • One number does not show clusters, curves, or uneven spread.

Squared correlation has a narrower interpretation

For a simple linear fit of fit on experience, \(r^2\approx 0.051\).

\[r^2=(0.227)^2\approx 0.051\]

Experience alone accounts for about 5% of the observed variation in fit under this simple linear summary.

  • This does not say experience causes 5% of fit, or that a richer model would perform similarly.

An observed relationship is not necessarily a causal effect

Applications with more experience tend to have higher fit, but the chart cannot identify why.

  • Applicants select jobs; jobs require different experience.
  • The resume-fit score may itself use experience-related inputs.
  • The workbook does not describe the scoring process or application selection.
  • Treat the pattern as a question for further study, not evidence that adding experience would change a score by a specific amount.

A scatterplot matrix checks several pairs

Change the observation to a job: compare application volume, mean experience, and mean fit across 403 jobs.

A correlation matrix is a compact summary

Summarize the same 403 jobs. Application volume and mean fit have a modest negative correlation (−0.24).

Correlation matrix: product subcategories

Heat map shows density in 2 dimensions

A two-dimensional count map reveals where many applications overlap in the scatterplot.

Numeric–category comparisons

For a numerical measure and a categorical field, compare full distributions before reducing each group to one number.

  • Use box plots to compare center and spread.
  • Use strips, violins, or small histograms to inspect shape.
  • Use dots or dumbbells when the question is about one summary per group.

Box plots compare experience across units

The median and spread of experience differ across business units.

Strip plots retain individual observations

A strip plot shows the range and concentration that group summaries compress.

A violin plot shows distribution shape

A mirrored density curve shows each group’s shape. The horizontal lines mark its median fit score.

A butterfly chart compares two distributions

Compare job–resume fit for male and female applicants. Each bar shows the share within its recorded gender category.

Small histograms compare the eight largest units

Repeated axes reveal whether the fit distribution changes across units.

Dumbbell charts are useful to compare numerical values across 2 categories

Each line connects the rate for applications with a referral to the rate for applications without one.

Small multiples of bars for categorical combinations

Within each unit, compare application counts across fields of study on the same scale.

Compare composition with 100% stacked bars

Platform Engineering applications are concentrated in technical fields. Customer-facing units draw a broader mix.

A colored frequency table supports exact lookup

Use the same unit × field-of-study counts to see both group size and composition.

Lifecycle of a data analysis project

Data acquisition

Provenance

  • Is the source reliable?
  • How was the data collected?
  • When was it collected?
  • What was the intent or motive?

Data understanding

  • What does the dataset contain?
  • What tables are available?
  • What does each row represent?
  • What data is missing?
  • Are there erroneous values?
  • Are the data types correct?

Data cleaning

  • What should we do with missing data?
  • Which data types need conversion?
  • What should we do with erroneous values?

Exploratory analysis

  • Explore variable distributions
  • Explore relationships in the data
  1. Ask a question / generate a hypothesis
  2. Create charts to answer the question
  3. Ask a new question based on the chart ↻ Repeat

Formal statistical modeling

Findings

Communication and dissemination