When you need cohort-scale clinical text — not one patient in a UI — you use
the FHIR Bulk Data Access ($export) operation: an async job that emits
NDJSON files of resources you then stream into OpenMed for batch
de-identification and NER. This skill sits before the OpenMed pipeline: it is
how the notes arrive.
Reach for it when the source is an EHR or FHIR data warehouse and the volume is
a population/group (thousands of patients), the workload is headless (no
clinician UI), and the goal is to batch-feed openmed.deidentify /
openmed.analyze_text. Triggers: "bulk export", "$export", "NDJSON", "Flat
FHIR", "cohort de-identification", "export all notes". For a single in-chart
patient with a UI, use scaffolding-smart-on-fhir instead.
GET [base]/$export — everything the client is authorized for.GET [base]/Group/[id]/$export — a defined cohort (most common).GET [base]/Patient/$export — all patients in scope.Bulk export uses SMART Backend Services auth (a system/*.read-scoped
client-credentials token via a signed JWT assertion), not an interactive launch.
# 1) Kickoff (async). Ask for clinical-note-bearing resource types.
curl -s -X GET \
'https://ehr.example/fhir/Group/cohort-42/$export?_type=DocumentReference,DiagnosticReport&_since=2024-01-01T00:00:00Z' \
-H 'Authorization: Bearer <backend-services-token>' \
-H 'Accept: application/fhir+json' \
-H 'Prefer: respond-async' -D -
# -> 202 Accepted
# Content-Location: https://ehr.example/fhir/bulkstatus/JOB123
# 2) Poll the status URL until complete
curl -s 'https://ehr.example/fhir/bulkstatus/JOB123' \
-H 'Authorization: Bearer <token>'
# 202 + X-Progress while running; 200 + a manifest JSON when done:
# { "transactionTime": "...", "request": "...", "requiresAccessToken": true,
# "output": [
# { "type": "DocumentReference",
# "url": "https://ehr.example/fhir/bulkfiles/dr-1.ndjson" },
# { "type": "DiagnosticReport",
# "url": "https://ehr.example/fhir/bulkfiles/dx-1.ndjson" } ] }
# 3) Download each NDJSON file (one FHIR resource per line)
curl -s 'https://ehr.example/fhir/bulkfiles/dr-1.ndjson' \
-H 'Authorization: Bearer <token>' -o dr-1.ndjson
Key headers/params: Prefer: respond-async (required to start the job),
Content-Location (the status/polling URL), _type (limit resource types),
_since (incremental export), _typeFilter (server-side resource filtering).
Delete the job when done: DELETE <status-url>.
NDJSON is one resource per line — stream it; do not load the whole file. Pull the
note text out of each DocumentReference/DiagnosticReport and run OpenMed
on-device, in batch:
import base64, json, openmed
def note_text(resource: dict) -> str | None:
# DocumentReference.content[].attachment.data (base64) or .url -> Binary
for content in resource.get("content", []):
att = content.get("attachment", {})
if att.get("data"):
return base64.b64decode(att["data"]).decode("utf-8", "replace")
# DiagnosticReport.presentedForm[].data
for form in resource.get("presentedForm", []):
if form.get("data"):
return base64.b64decode(form["data"]).decode("utf-8", "replace")
return None
with open("dr-1.ndjson", "r", encoding="utf-8") as fh:
for line in fh: # streaming, line by line
resource = json.loads(line)
text = note_text(resource)
if not text:
continue
# De-identify every note before anything downstream sees it
deid = openmed.deidentify(text, method="replace", policy="hipaa_safe_harbor")
# Then NER on the de-identified text
entities = openmed.analyze_text(
deid.text, model_name="disease_detection_superclinical")
# ... persist de-identified text + spans; never persist raw PHI
For large cohorts, parallelise across files (each NDJSON file is independent) and reuse a single OpenMed model loader across notes to avoid reloading weights.
system/DocumentReference.read, etc.).$export at the right level with _type (and _since for
incrementals) + Prefer: respond-async.Content-Location until 200; read the manifest output[].requiresAccessToken).openmed.deidentify →
openmed.analyze_text.exporting-to-fhir,
assembling-fhir-bundles).DELETE the bulk job to free server storage.openmed.deidentify is the primary hand-off. De-identify first; treat
every exported note as PHI until it has been through the de-id pass.analyze_text → exporting-to-fhir →
to_bundle; write back only if your governance allows.Content-Location is
success; poll with backoff and honour Retry-After/X-Progress.json.load a whole
file. Parallelise per file, not per line.requiresAccessToken. If the manifest says so, send the bearer token when
downloading the NDJSON files too.openmed.eval
leakage gates (evaluating-with-leakage-gates), not F1 alone.Binary
reference, or RTF/HTML in presentedForm. Normalise to plain text before
OpenMed; for scanned PDFs use OpenMed's document/OCR intake.DELETE the
status URL when finished.$export operation: https://hl7.org/fhir/uv/bulkdata/export.html