Research-ready imaging datasets
From your archive to a research-ready dataset
One pipeline extracts, de-identifies, normalizes and delivers your imaging archive as a queryable, SNOMED-coded dataset. Your PACS keeps running. You own the result.
Every archive we open has these
What real archives look like before they become datasets
We have run the pipeline on archives from hospitals, imaging networks and data companies, across every modality and vendor. The same problems show up every time, and every one of them breaks a cohort query, a de-identification pass, or both.
No body part at all
The single most common gap: series with no BodyPartExamined value, often the majority of an archive. The image shows a lumbar spine; the header says nothing.
The label is wrong, or says “Other”
A series described as “Abdomen” whose pixels show a knee. Thousands filed under “Other”. Header-only tools file them wrong forever. A pixel check does not.
Modality contradicts the object
Headers say “Other” or “Secondary Capture” while the pixels are a SPECT or a CT key-image set. Filter a cohort on modality and they silently drop out.
Names in free-text fields
Patient names typed into series descriptions and protocol names. Tag-list de-identification never looks there; a pipeline that reads every field does.
Descriptions that say nothing, in any language
“AX ST”, “One Shot”, “Series 1”, “Key Images”, “AXI DIFUSAO”, “MEDIASTINO”. Unqueryable until every one is rebuilt into a standard description.
Damage from earlier anonymization
Name placeholders written into the Modality tag. GE private tags stripped, and with them the T1/T2/FLAIR weighting the study needs. We detect it, restore what can be restored, and flag the rest.
Why de-identification fails at scale
Tag scrubbers and one-off scripts work on a hundred studies. They break on a million.
✕ Tag-only de-identification leaves burned-in PHI in the pixels and names in free-text fields
✓ Pixel-level OCR redaction plus every text field read, validated against the image itself
✕ Blanket stripping destroys the private tags that make data clinically usable: on GE MR, that is where T1, T2 and FLAIR weighting live
✓ Selective handling: private PHI scrubbed, clinical vendor tags preserved
✕ Per-study manual review can’t cover millions of studies
✓ Review batched by finding: one approval fixes thousands of studies
One inventory, any number of datasets
Before anything is de-identified, the whole corpus is read, normalized and indexed. The result is an inventory you can query like a database: every study described the same way, coded in SNOMED CT, with laterality, contrast and anatomy resolved. Define a cohort, get a dataset. Define another, get another. The pipeline runs once; the datasets are generated from it, and new studies can flow in monthly.
EXAMPLE — CREATING ONE COHORT
QUERY
Modality = CT · Region = Chest · Contrast = none · Adults · 2018–2025
INVENTORY ANSWERS
9,840 studies · 5 sites · 3 vendors · Body part resolved: 100% (was 41%)
DATASET GENERATED
De-identified, dates shifted, longitudinal links kept · Validation workbook + QA sample
NEXT COHORT
Same inventory, new query · No re-extraction
Three levels of delivery
Each level builds on the one before. Choose per cohort, not per archive.
Normalized
• Metadata normalized across study, series and body part
• SNOMED CT coding at study and series level, validated by sampling
• Clinical vendor tags preserved
• Queryable inventory over the whole corpus
Identified data, for migration and internal use
De-identified
• Everything in Normalized, plus:
• PHI removed at tag and pixel level (burned-in text, OCR)
• UIDs remapped, dates shifted, longitudinal links preserved
• Private vendor tags scrubbed, head and neck defaced
• Validation workbook and QA sampling report
A research-ready dataset for internal research and AI training
Coded and linked
• Everything in De-identified, plus:
• SNOMED CT coding verified on every instance, with laterality
• Reports and structured reports linked, free text de-identified
• Provenance pack and expert-determination support
For external collaboration and licensing, on your terms
Inside your network or in your own tenant
De-identification can run on a connector inside your network, so PHI is removed before anything leaves the building, or in a dedicated single-client tenant in the EU or US with full compute for pixel-level redaction. Nothing installs on the PACS; reads are throttled. Every removal and correction is logged and approved by your team.
SOC 2 Type II
ISO 27001
HIPAA
GDPR
Find out what’s actually in your archive
Most engagements start small: a Data Quality Report on 10,000 studies from your archive, any modality, any vendor. Full DICOM-hierarchy analysis, burned-in PHI exposure by modality, de-identification feasibility on your actual data, and sample corrected files, delivered in 4–6 weeks with a scoped plan and per-study pricing for the full archive. It begins with a 30-minute intro call.

