Profile a CSV: nulls, types and top values
Tested by BSVWrite profile.csv and profile.md for any CSV. With no path, the script creates a labeled synthetic sample first.
- Input
- Optional path to a CSV. Default: a short labeled synthetic sample written by the script.任意:CSVのパス。省略時はスクリプトが短いラベル付き合成サンプルを書きます。
- Output
- profile.csv (one row per column) and profile.md (human-readable summary).profile.csv(列ごと1行)と profile.md(人が読む要約)。
- Prerequisites
- Python 3 stdlib only (csv).Python 3 標準ライブラリのみ(csv)。
- Steps
-
- Save the script.
- Run: python csv_profile.py or python csv_profile.py your.csv
- Open profile.md; confirm null counts before you trust any chart.
- スクリプトを保存します。
- 実行: python csv_profile.py または python csv_profile.py your.csv
- profile.md を開き、グラフを信じる前に欠損件数を確認します。
- Expected result
- A profile for every column. Synthetic samples stay labeled example-row — replace with your own file for real work.全列のプロフィール。合成サンプルは example-row のままです。本番では自分のファイルに差し替えてください。
- Next step
- Flatten a JSONL export next, or pull a public time series once local files look clean.次は JSONL の平坦化か、ローカルがきれいに見えたら公開時系列の取得へ。 Open
Code (BSV original, MIT licence)
Download .py · csv_profile.py
#!/usr/bin/env python3
"""BSV recipe: profile a CSV — column types, nulls, distinct counts, sample values.
Input : path to a CSV (default: writes and profiles a labeled synthetic sample).
Output: profile.csv + profile.md in the working directory.
Original BSV code, MIT. Synthetic sample rows are labeled examples, not customer or production data.
"""
import csv, os, sys
from collections import Counter
SAMPLE = """id,city,temp_c,note
1,Tokyo,22.1,example-row
2,Osaka,,example-row
3,Tokyo,19.4,example-row
4,Nagoya,21.0,example-row
5,Osaka,18.7,example-row
"""
def is_float(s):
try:
float(s); return True
except Exception:
return False
def main(path=None):
if not path:
path = "sample_labeled.csv"
open(path, "w", encoding="utf-8").write(SAMPLE)
print("wrote labeled synthetic sample:", path)
rows = list(csv.DictReader(open(path, encoding="utf-8", newline="")))
if not rows:
raise SystemExit("empty csv")
cols = list(rows[0].keys())
out_rows = []
lines = [f"# CSV profile for `{path}`", "", f"Rows: {len(rows)} Columns: {len(cols)}", "",
"Synthetic or practice files should stay labeled as such; do not treat sample numbers as live metrics.", ""]
for c in cols:
vals = [r.get(c, "") for r in rows]
empty = sum(1 for v in vals if v is None or str(v).strip() == "")
filled = [str(v).strip() for v in vals if v is not None and str(v).strip() != ""]
numeric = all(is_float(v) for v in filled) if filled else False
distinct = len(set(filled))
top = Counter(filled).most_common(3)
samples = ", ".join(repr(v) for v, _ in top) if top else ""
out_rows.append({"column": c, "non_null": len(filled), "nulls": empty,
"distinct": distinct, "numeric": str(numeric).lower(), "top_values": samples})
lines.append(f"## {c}")
lines.append(f"- non-null: {len(filled)} / nulls: {empty} / distinct: {distinct} / numeric: {numeric}")
lines.append(f"- top: {samples}")
lines.append("")
with open("profile.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=["column", "non_null", "nulls", "distinct", "numeric", "top_values"])
w.writeheader(); w.writerows(out_rows)
open("profile.md", "w", encoding="utf-8").write("\n".join(lines) + "\n")
print(f"profiled {len(rows)} rows × {len(cols)} cols -> profile.csv, profile.md")
if __name__ == "__main__":
main(*sys.argv[1:])
Output of BSV's own test run
Run on 2026-10-07 02:01 JST · Python 3.13.5
wrote labeled synthetic sample: sample_labeled.csv
profiled 5 rows × 4 cols -> profile.csv, profile.md