Skip to main content
DATASET DOCTOR

Open-source dataset diagnostics

Know your dataset
before you train
your model.

Catch data quality problems before they become model problems. Dataset Doctor helps surface missing values, duplicates, imbalance, suspicious features and other common dataset issues in a clear report.

A design concept — an open-source Python toolkit in the making. Every figure here is illustrative.

Open source
Python-friendly
Built for ML workflows
Dataset ReportSample output
12458 rows · 31 columnsSample analysis: 1.2s
FindingSeverity
  • Missing Values4 columnsWarning
  • Duplicate Rows218Information
  • Class Imbalance87 / 13Warning
  • Potential Leakagecustomer_statusCritical
Health Score
68 / 100Needs attention

Illustrative sample report · report.html

Designed for modern machine-learning workflows.

Why it matters

Model problems often begin before training starts.

Most teams discover dataset problems too late — after hours spent tuning architectures and hyperparameters that were never the bottleneck. By then the fix belongs upstream, and the experiments have to be re-run.

“A model trained on a sick dataset learns the sickness.”

Symptom indexS-01 — S-06
  1. 01Missing dataSilent NaN drift in 4 columns
  2. 02Duplicate observations218 rows recorded twice
  3. 03Class imbalance87 / 13 split between priors
  4. 04Suspicious featuresTarget proxies in disguise
  5. 05Incorrect data typesNumerics parsed as strings
  6. 06Data leakageAnswer keys in the training frame

The checklist

Everything worth checking
before model.fit()

Eight passes over the schema, the statistics and the relationships between them — the quiet checks that decide whether a model will generalize.

Pass 01 · Completeness

Missing Value Analysis

Null patterns grouped by column, dtype and row segment — not just one sad total count.

Filled squares mark present values, red outlines mark nulls — 4 of 31 columns carry nulls · 0.8% of cells

Pass 02 · Uniqueness

dup

Duplicate Detection

Exact and near-duplicate rows caught by fingerprint hashing, across every column pair.

Row bars preview record fingerprints; matching rows highlight amber — 218 rows collapse into 109 unique records

Pass 03

Class Distribution

Majority and minority priors, surfaced before they skew metrics.

Pass 04

Outlier Detection

Values drifting far outside their column's expected range.

Pass 05

Data Type Validation

Numerics stored as strings, dates stored as anything but.

Pass 06

Leakage Warnings

Features holding answers they should not have access to.

Pass 07 · Relationships

Correlation Analysis

Which features move together — and which ones happen to move with the target.

Illustrative correlation mapExample finding: field_12 and field_27 have ρ 0.94 — review for leakage

Overview

Dataset Summary

The factual shape of your data, before any modelling decision.

Rows
12,458
Columns
31
Numeric
22
Categorical
9
Memory
14.2 MB

Dataset Health

One score. A clearer starting point.

Dataset Health summarizes common quality signals to help identify which areas deserve attention before training.

68/ 100Overall score
Needs attention
Diagnostic categoriesScore

Completeness

nulls in 4 of 31 columns

92

Consistency

mixed encodings in region

81

Class Balance

87 / 13 class prior

44

Duplicate Quality

218 duplicate rows

76

Leakage Risk

customer_status · ρ 0.94

55
  • 80–100 strong
  • 50–79 watch
  • 0–49 act now

Concept preview — this page is a product concept. Scores and findings shown here are illustrative, not computed from real data.

Sample report

Problems, prioritized.

Every finding ships with evidence and a suggested next action — sorted by severity, so the first hour goes to the biggest risk.

Selected findings — the full report lists every check with its evidence.

LEAK-001

Potential target leakage

Critical

The feature "customer_status" has a suspiciously strong relationship with the target.

Evidence

customer_statusmedian feature

Suggested action

Review whether this column would actually be available at prediction time.

BAL-004

Class imbalance detected

Warning

Training on this split will bias the model toward the majority class unless it is weighted or resampled.

Evidence

Class A · 87%Class B · 13%

Suggested action

Weight the classes or resample the training split before fitting; report precision and recall per class instead of plain accuracy.

DUP-012

Duplicate observations

Information

218 duplicated rows detected.

Evidence

218

duplicated rows · 1.75% of dataset

Suggested action

Keep one copy of each duplicated record and regenerate the split before training — or trace why duplicates enter the pipeline upstream.

How it works

From dataset to insight
in three steps.

  1. 01csv · parquet

    Upload

    CSV or Parquet dataset.

  2. 02Illustrative checks

    Inspect

    Automated quality and ML-readiness checks.

  3. 03report.html

    Improve

    Fix issues before training.

Workflow integration

Fits into the workflow you already use.

Run one command in CI, or import the analyzer beside your notebook. Either way the report arrives as a single self-contained HTML file.

Command and API shown are illustrative — the interface is still being shaped.

Sample session
$ dataset-doctor inspect dataset.csvAnalyzing dataset...✓ 12458 rows✓ 31 features! 3 warnings✕ 2 critical issuesDataset health: 68 / 100✓ report.html generated

Open source

Transparent by design.

Dataset Doctor is envisioned as an open-source toolkit that developers can inspect, extend and integrate into their existing workflows.

Open Source

Read every check, audit every heuristic, ship your own fork.

MIT license

Python First

A proposed interface for notebooks, scripts and pipelines.

Proposed Python interface

Extensible Checks

A proposed catalogue of checks with room for custom validation passes.

Custom checks concept

Inspect first.Train second.

Catch obvious dataset problems before spending time optimizing the wrong model.

Product concept · sample data throughout