Pass 01 · Completeness
Missing Value Analysis
Null patterns grouped by column, dtype and row segment — not just one sad total count.
Filled squares mark present values, red outlines mark nulls — 4 of 31 columns carry nulls · 0.8% of cells
Open-source dataset diagnostics
Catch data quality problems before they become model problems. Dataset Doctor helps surface missing values, duplicates, imbalance, suspicious features and other common dataset issues in a clear report.
A design concept — an open-source Python toolkit in the making. Every figure here is illustrative.
Illustrative sample report · report.html
Designed for modern machine-learning workflows.
Why it matters
Most teams discover dataset problems too late — after hours spent tuning architectures and hyperparameters that were never the bottleneck. By then the fix belongs upstream, and the experiments have to be re-run.
“A model trained on a sick dataset learns the sickness.”
The checklist
Eight passes over the schema, the statistics and the relationships between them — the quiet checks that decide whether a model will generalize.
Pass 01 · Completeness
Null patterns grouped by column, dtype and row segment — not just one sad total count.
Filled squares mark present values, red outlines mark nulls — 4 of 31 columns carry nulls · 0.8% of cells
Pass 02 · Uniqueness
dupExact and near-duplicate rows caught by fingerprint hashing, across every column pair.
Row bars preview record fingerprints; matching rows highlight amber — 218 rows collapse into 109 unique records
Pass 03
Majority and minority priors, surfaced before they skew metrics.
Pass 04
Values drifting far outside their column's expected range.
Pass 05
Numerics stored as strings, dates stored as anything but.
Pass 06
Features holding answers they should not have access to.
Pass 07 · Relationships
Which features move together — and which ones happen to move with the target.
Overview
The factual shape of your data, before any modelling decision.
Dataset Health
Dataset Health summarizes common quality signals to help identify which areas deserve attention before training.
Completeness
nulls in 4 of 31 columns
Consistency
mixed encodings in region
Class Balance
87 / 13 class prior
Duplicate Quality
218 duplicate rows
Leakage Risk
customer_status · ρ 0.94
Concept preview — this page is a product concept. Scores and findings shown here are illustrative, not computed from real data.
Sample report
Every finding ships with evidence and a suggested next action — sorted by severity, so the first hour goes to the biggest risk.
Selected findings — the full report lists every check with its evidence.
The feature "customer_status" has a suspiciously strong relationship with the target.
Evidence
Suggested action
Review whether this column would actually be available at prediction time.
Training on this split will bias the model toward the majority class unless it is weighted or resampled.
Evidence
Suggested action
Weight the classes or resample the training split before fitting; report precision and recall per class instead of plain accuracy.
218 duplicated rows detected.
Evidence
218
Suggested action
Keep one copy of each duplicated record and regenerate the split before training — or trace why duplicates enter the pipeline upstream.
How it works
CSV or Parquet dataset.
Automated quality and ML-readiness checks.
Fix issues before training.
Workflow integration
Run one command in CI, or import the analyzer beside your notebook. Either way the report arrives as a single self-contained HTML file.
Command and API shown are illustrative — the interface is still being shaped.
$ dataset-doctor inspect dataset.csvAnalyzing dataset...✓ 12458 rows✓ 31 features! 3 warnings✕ 2 critical issuesDataset health: 68 / 100✓ report.html generatedOpen source
Dataset Doctor is envisioned as an open-source toolkit that developers can inspect, extend and integrate into their existing workflows.
Read every check, audit every heuristic, ship your own fork.
MIT license
A proposed interface for notebooks, scripts and pipelines.
Proposed Python interface
A proposed catalogue of checks with room for custom validation passes.
Custom checks concept
Catch obvious dataset problems before spending time optimizing the wrong model.
Product concept · sample data throughout