A healthcare robotics company needed to deliver de-identified DICOM imaging data, CT scans, X-rays, and angiograms, to providers and data scientists without exposing protected health information. We built an event-driven AWS pipeline that automates ingestion, PHI masking, and secure querying end to end.
Talk to UsRemote-operated medical robotics treating patients across geographic barriers, built on medical imaging data that had to reach providers and data scientists PHI-free.
End-to-end DICOM ingestion, PHI masking, and secure querying pipeline, from raw S3 upload to ML-ready, de-identified imaging data.
AWS Lambda, S3, SQS, DynamoDB, Fargate, Python, SQL, Databricks, V7 Labs.
Delivered
The pipeline replaced fully manual DICOM handling with automated ingestion and masking. Processing time dropped from weeks to a single day, and the Databricks integration brought over 40% improvement in data processing and analytics capability on top of it.
Success rate in DICOM file management
Efficiency improvement across processing workflows
PHI had to come off every image, file handling had to stop being manual, subsets had to be queryable at scale, and none of it could happen without first untangling how the raw uploads were structured.
Every DICOM upload arrived as a full folder, mixing medical and non-medical files together. Isolating the actual DICOM files had to happen before any de-identification work could start, with no prior team depth in DICOM file parameters to draw on.
Every image carried protected health information that had to be removed before the underlying imaging data could be used to train models.
No automated processing or structured storage existed for uploaded files, so every file moved through the pipeline by hand.
The client needed to pull specific file subsets against research-specific criteria at scale, with strong authentication protecting the entire querying process.
An AWS-native pipeline handles ingestion and masking, with Databricks and V7 Labs layered on for modeling, annotation, and visualization.
DICOM files land in an S3 dirty bucket the moment they are uploaded. That upload event triggers Lambda functions that process and PHI-mask each file, then write the de-identified result to a separate clean bucket. SQS sits between the upload event and the Lambda processing, decoupling the two so the pipeline scales across multiple concurrent queries without bottlenecking.
Patient-identifiable fields in each DICOM file convert to unique, non-reversible hash keys. The underlying imaging content, the CT, X-ray, and angiogram data, stays intact and usable for ML training, while the original patient record cannot be reconstructed from the hashed output.
Once files are de-identified, they feed into Databricks for big data processing, data engineering, and ML modeling. V7 Labs handles data annotation and visualization on the same processed imaging data, so data scientists work from a single de-identified source instead of separate raw exports.
Every login gets a personalized view showing query history and applied scripts, so users can revisit or refine prior searches without rebuilding them. Multi-factor authentication, encryption at rest and in transit, and role-based access controls run across the entire querying process.
Need PHI masked out before it reaches your ML pipeline? Talk to us about what an automated DICOM processing pipeline looks like for your team.
Partner with us to design, build, and scale digital solutions that drive better outcomes.
Global Tech Teams LLC, 525 Washington Blvd, Industrious at Newport Tower, Jersey City, NJ 07310, United States.
Let’s discuss your goals, workflows, and next steps in a focused consultation call.