Skills Data Science Redacting Free-Text Clinical Data

Redacting Free-Text Clinical Data

v20260803
deidentify-a-dataset
This tool de-identifies sensitive free-text columns (e.g., names, dates, identifiers) in local structured datasets (CSV, JSONL, Parquet) using advanced OpenMed models. It ensures strict privacy by generating a completely redacted dataset and a separate PHI-free aggregate summary, allowing for secure data sharing and analysis without compromising source privacy or logging sensitive values.
Get Skill
362 downloads
Overview

De-identify a dataset

Keep the source local, name the free-text columns explicitly, and write to a different destination. Never infer columns or print source and redacted cell values.

Procedure

  1. Confirm that the input is CSV, JSONL/NDJSON, or Parquet.
  2. Confirm which columns contain free text. Do not scan or log values to guess.
  3. Choose a policy and language. Prefer strict_no_leak when recall is the governing safety requirement.
  4. Write to a new path; never overwrite the input.
  5. Inspect only result.summary, which contains aggregate counts and rates.
  6. Validate recall and residual leakage on representative synthetic or approved evaluation fixtures before releasing the output.

Runnable synthetic example

Install the model runtime first with python -m pip install "openmed[hf]".

import csv
from pathlib import Path

from openmed import redact_dataset

source = Path("synthetic-notes.csv")
destination = Path("synthetic-notes.redacted.csv")

with source.open("w", newline="", encoding="utf-8") as handle:
    writer = csv.DictWriter(handle, fieldnames=["record_id", "note"])
    writer.writeheader()
    writer.writerows(
        [
            {
                "record_id": "SYNTH-001",
                "note": (
                    "Taylor Example called 212-555-0198 about a "
                    "metformin refill."
                ),
            },
            {
                "record_id": "SYNTH-002",
                "note": (
                    "Send the synthetic follow-up to "
                    "demo.patient@example.test."
                ),
            },
        ]
    )

result = redact_dataset(
    source,
    text_columns=["note"],
    output_path=destination,
    policy="strict_no_leak",
    lang="en",
)

print(result.output_path)
print(result.summary.to_dict())  # Aggregate counts only; no cell contents.

Use the equivalent CLI for an existing dataset:

openmed redact-dataset notes.csv \
  --text-columns note,comment \
  --policy strict_no_leak \
  --output notes.redacted.csv

Safety checks

  • Keep model inference and files on infrastructure the user controls.
  • Do not print input rows, detected entity surfaces, reversible mappings, or exception payloads that may contain source text.
  • Keep source and output paths separate and access-controlled.
  • Treat the aggregate summary as evidence, not as proof of compliance.
  • Never commit real clinical data or restricted evaluation corpora.

Repository example

Read and run the offline dataset walkthrough when you need a bundled synthetic fixture and first-run download controls.

Info
Category Data Science
Name deidentify-a-dataset
Version v20260803
Size 2.97KB
Updated At 2026-08-04
Language