Research-ready imaging datasets

From your archive to a research-ready dataset

One pipeline extracts, de-identifies, normalizes and delivers your imaging archive as a queryable, SNOMED-coded dataset. Your PACS keeps running. You own the result.

Datamonk research inventory and cohort builder
Datamonk research inventory and cohort builder

Every archive we open has these

What real archives look like before they become datasets

We have run the pipeline on archives from hospitals, imaging networks and data companies, across every modality and vendor. The same problems show up every time, and every one of them breaks a cohort query, a de-identification pass, or both.

No body part at all

The single most common gap: series with no BodyPartExamined value, often the majority of an archive. The image shows a lumbar spine; the header says nothing.

The label is wrong, or says “Other”

A series described as “Abdomen” whose pixels show a knee. Thousands filed under “Other”. Header-only tools file them wrong forever. A pixel check does not.

Modality contradicts the object

Headers say “Other” or “Secondary Capture” while the pixels are a SPECT or a CT key-image set. Filter a cohort on modality and they silently drop out.

Names in free-text fields

Patient names typed into series descriptions and protocol names. Tag-list de-identification never looks there; a pipeline that reads every field does.

Descriptions that say nothing, in any language

“AX ST”, “One Shot”, “Series 1”, “Key Images”, “AXI DIFUSAO”, “MEDIASTINO”. Unqueryable until every one is rebuilt into a standard description.

Damage from earlier anonymization

Name placeholders written into the Modality tag. GE private tags stripped, and with them the T1/T2/FLAIR weighting the study needs. We detect it, restore what can be restored, and flag the rest.

Why de-identification fails at scale

Tag scrubbers and one-off scripts work on a hundred studies. They break on a million.

✕ Tag-only de-identification leaves burned-in PHI in the pixels and names in free-text fields

✓ Pixel-level OCR redaction plus every text field read, validated against the image itself

✕ Blanket stripping destroys the private tags that make data clinically usable: on GE MR, that is where T1, T2 and FLAIR weighting live

✓ Selective handling: private PHI scrubbed, clinical vendor tags preserved

✕ Per-study manual review can’t cover millions of studies

✓ Review batched by finding: one approval fixes thousands of studies

One inventory, any number of datasets

Before anything is de-identified, the whole corpus is read, normalized and indexed. The result is an inventory you can query like a database: every study described the same way, coded in SNOMED CT, with laterality, contrast and anatomy resolved. Define a cohort, get a dataset. Define another, get another. The pipeline runs once; the datasets are generated from it, and new studies can flow in monthly.

EXAMPLE — CREATING ONE COHORT

QUERY

Modality = CT · Region = Chest · Contrast = none · Adults · 2018–2025

INVENTORY ANSWERS

9,840 studies · 5 sites · 3 vendors · Body part resolved: 100% (was 41%)

DATASET GENERATED

De-identified, dates shifted, longitudinal links kept · Validation workbook + QA sample

NEXT COHORT

Same inventory, new query · No re-extraction

Three levels of delivery

Each level builds on the one before. Choose per cohort, not per archive.

Normalized

• Metadata normalized across study, series and body part
• SNOMED CT coding at study and series level, validated by sampling
• Clinical vendor tags preserved
• Queryable inventory over the whole corpus

Identified data, for migration and internal use

De-identified

• Everything in Normalized, plus:
• PHI removed at tag and pixel level (burned-in text, OCR)
• UIDs remapped, dates shifted, longitudinal links preserved
• Private vendor tags scrubbed, head and neck defaced
• Validation workbook and QA sampling report

A research-ready dataset for internal research and AI training

Coded and linked

• Everything in De-identified, plus:
• SNOMED CT coding verified on every instance, with laterality
• Reports and structured reports linked, free text de-identified
• Provenance pack and expert-determination support

For external collaboration and licensing, on your terms

Inside your network or in your own tenant

De-identification can run on a connector inside your network, so PHI is removed before anything leaves the building, or in a dedicated single-client tenant in the EU or US with full compute for pixel-level redaction. Nothing installs on the PACS; reads are throttled. Every removal and correction is logged and approved by your team.

SOC 2 Type II

ISO 27001

HIPAA

GDPR

Find out what’s actually in your archive

Most engagements start small: a Data Quality Report on 10,000 studies from your archive, any modality, any vendor. Full DICOM-hierarchy analysis, burned-in PHI exposure by modality, de-identification feasibility on your actual data, and sample corrected files, delivered in 4–6 weeks with a scoped plan and per-study pricing for the full archive. It begins with a 30-minute intro call.

Frequently asked questions

Do we have to migrate to get a research dataset?

How do you handle burned-in text in the images?

Who decides what gets removed or changed?

Can we license the dataset to third parties?

What does it cost?