Why do we clean data?
Measurements can be missing, entered incorrectly, repeated or extreme. Before using averages and models, inspect which observations you can trust. This beginner lesson uses a NumPy-enabled Python environment.
Recognise missing data
A floating-point array can contain `np.nan`. Missing values are different from zero. Ordinary reductions may propagate NaN.
import numpy as np
scores = np.array([76., np.nan, 82., 90.])
print(np.isnan(scores))
print(np.mean(scores)) # nan
print(np.nanmean(scores)) # 82.666...Remove missing entries
Boolean masks let you keep observed values. For numeric arrays use `~np.isnan(a)`; NaNs should not be compared with `==`.
a = np.array([10., np.nan, 20., 30.])
clean = a[~np.isnan(a)]
print(clean) # [10. 20. 30.]Replace missing data
Choose a replacement that makes sense for your problem. Mean imputation is simple but can hide uncertainty. Work on a copy to preserve raw data.
a = np.array([50., np.nan, 70., 80.])
filled = a.copy()
filled[np.isnan(filled)] = np.nanmean(filled)
print(filled)Clean invalid values
A negative price or a score above 100 may be invalid. Filter using domain rules rather than removing every unexpected value blindly.
marks = np.array([45, 105, -4, 89, 72])
valid = marks[(marks >= 0) & (marks <= 100)]
print(valid) # [45 89 72]Remove duplicate values
`np.unique` returns unique values sorted by default. Use care: repeated measurements or repeated scores are not necessarily duplicate records.
a = np.array([4, 2, 4, 1, 2])
print(np.unique(a)) # [1 2 4]
values, counts = np.unique(a, return_counts=True)
print(values, counts)Find outliers with IQR
Use quartiles to flag unusually distant values. An outlier flag is not proof of an error. Review flagged cases before deleting.
a = np.array([10., 11., 12., 13., 14., 100.])
q1, q3 = np.percentile(a, [25, 75])
iqr = q3 - q1
low, high = q1 - 1.5 * iqr, q3 + 1.5 * iqr
print("Possible outliers:", a[(a < low) | (a > high)])Normalise safely
Min-max scaling maps values into the 0–1 interval. Protect against a zero range when all values are equal.
a = np.array([40., 50., 80.])
span = np.ptp(a)
scaled = (a - a.min()) / span if span else np.zeros_like(a)
print(scaled)Clean a two-dimensional table
Keep rows and columns aligned when cleaning tabular data. Filter rows using a row-wise Boolean mask.
marks = np.array([[70., 80.], [np.nan, 60.], [90., 75.]])
complete_rows = marks[~np.isnan(marks).any(axis=1)]
print(complete_rows)Mini project: Student score audit
Preserve original input, count missing data, reject invalid scores, fill legitimate missing values with subject medians and create a summary. For a real dataset, record why each value was corrected.
import numpy as np
raw = np.array([[80., 75., np.nan],
[65., 110., 72.],
[90., 85., 88.],
[np.nan, 70., 60.]])
clean = raw.copy()
invalid = (clean < 0) | (clean > 100)
print("Invalid entries:", invalid.sum())
clean[invalid] = np.nan
print("Missing after validation:", np.isnan(clean).sum())
for col in range(clean.shape[1]):
median = np.nanmedian(clean[:, col])
if np.isnan(median):
raise ValueError("A subject has no valid scores")
missing = np.isnan(clean[:, col])
clean[missing, col] = median
print("Cleaned data:\n", clean)
print("Student averages:", clean.mean(axis=1))
print("Subject averages:", clean.mean(axis=0))Practice exercises
- Count missing values with
np.isnan(). - Replace
np.nanusing a column-wise median. - Remove invalid ages outside a realistic range, using a clearly stated rule.
- Inspect repeated values with
np.unique(return_counts=True). - Find IQR outlier candidates, then explain whether they are genuine errors.
Challenge: Adapt the student score audit to four subjects and calculate the average after cleaning.
Continue from 0 to infinity
Previous: NumPy Array Operations. Suggested next lesson: Pandas and DataFrames for cleaning labelled, mixed-type tables.
Python editor (choose an environment with NumPy installed to run this lesson).