DICOM Processing Pipeline Cuts Turnaround From Weeks to a Single Day

A healthcare robotics company needed to deliver de-identified DICOM imaging data, CT scans, X-rays, and angiograms, to providers and data scientists without exposing protected health information. We built an event-driven AWS pipeline that automates ingestion, PHI masking, and secure querying end to end.

Talk to Us
Customer Focus

Remote-operated medical robotics treating patients across geographic barriers, built on medical imaging data that had to reach providers and data scientists PHI-free.

Scope

End-to-end DICOM ingestion, PHI masking, and secure querying pipeline, from raw S3 upload to ML-ready, de-identified imaging data.

Stack

AWS Lambda, S3, SQS, DynamoDB, Fargate, Python, SQL, Databricks, V7 Labs.

Status

Delivered

Outcomes

Weeks of Manual Review Down to a Single Day

The pipeline replaced fully manual DICOM handling with automated ingestion and masking. Processing time dropped from weeks to a single day, and the Databricks integration brought over 40% improvement in data processing and analytics capability on top of it.

99%

Success rate in DICOM file management

30%

Efficiency improvement across processing workflows

The Problem

Four Problems Before Any ML Training Could Start

PHI had to come off every image, file handling had to stop being manual, subsets had to be queryable at scale, and none of it could happen without first untangling how the raw uploads were structured.

01
No clean file boundary

Every DICOM upload arrived as a full folder, mixing medical and non-medical files together. Isolating the actual DICOM files had to happen before any de-identification work could start, with no prior team depth in DICOM file parameters to draw on.

02
Unmasked PHI blocked ML use

Every image carried protected health information that had to be removed before the underlying imaging data could be used to train models.

03
Fully manual file handling

No automated processing or structured storage existed for uploaded files, so every file moved through the pipeline by hand.

04
No secure way to query subsets

The client needed to pull specific file subsets against research-specific criteria at scale, with strong authentication protecting the entire querying process.

The Tech Stack

An AWS-native pipeline handles ingestion and masking, with Databricks and V7 Labs layered on for modeling, annotation, and visualization.

  • AWS Lambda
  • Amazon S3
  • Amazon SQS
  • DynamoDB
  • AWS Fargate (ECS)
  • Python
  • SQL
  • Databricks
  • V7 Labs
What We Built

An Automated Path From Raw Upload to ML-Ready Data

Event-Driven Ingestion on AWS

DICOM files land in an S3 dirty bucket the moment they are uploaded. That upload event triggers Lambda functions that process and PHI-mask each file, then write the de-identified result to a separate clean bucket. SQS sits between the upload event and the Lambda processing, decoupling the two so the pipeline scales across multiple concurrent queries without bottlenecking.

  • S3 dirty bucket captures every raw DICOM upload as it arrives
  • Lambda functions trigger automatically on upload to process and mask each file
  • SQS decouples upload events from processing so concurrent queries scale
  • De-identified output writes to a separate clean bucket, isolated from raw files

Processing medical imaging data at scale?

Need PHI masked out before it reaches your ML pipeline? Talk to us about what an automated DICOM processing pipeline looks like for your team.

Talk to Us

Let’s #Transform Healthcare,# Together.

Partner with us to design, build, and scale digital solutions that drive better outcomes.

Location

Global Tech Teams LLC, 525 Washington Blvd, Industrious at Newport Tower, Jersey City, NJ 07310, United States.

Contact

+1 408 786 5974
contact@mindbowser.com
BOOK A QUICK CONSULTATION

Have a Healthcare Project in Mind?

Let’s discuss your goals, workflows, and next steps in a focused consultation call.

Calendar icon Schedule a Call

Contact form