This post was originally published on this site

The intersection of medicine and AI has led to remarkable innovations. However, developers now face the thorny challenge of building robust medical AI tools that have been tested and evaluated on diverse, real-world patient data while also protecting patient privacy. At Google Cloud, our approach combines strategic collaboration with Confidential Computing.

To help protect both patient privacy and AI models during validation, we’re collaborating with MLCommons through the MedPerf initiative. First announced at Google Cloud Next earlier this year, this partnership uses Confidential Computing to establish a secure clean room for benchmarking AI models in real-world settings.

The challenge: Evaluating AI without seeing the data

MLCommons, a global community with over 125 members across tech and academia, launched MedPerf in 2023 to standardize the evaluation of medical AI. MedPerf, an open-source platform for benchmarking AI models, has advanced clinical research using federated evaluation to test models.

By using Google Cloud Confidential Space, proprietary AI models can be evaluated inside hardware-isolated Trusted Execution Environments (TEEs). This special virtual machine encrypts memory in-use and hardens the operating system, so none of the parties — the hospital or research institution, other participants, or Google — can see model code or patient data while it’s evaluated.

Medical AI benchmarking is compute-heavy, so the Confidential VM extends beyond the CPU to the GPU. To protect model weights and patient data even during GPU-accelerated inference, MedPerf runs on Google Cloud’s A3 machine series with NVIDIA H100 GPUs, which pairs Intel TDX technology on the CPU with NVIDIA Confidential Computing on the GPU. 

Before any patient data is released into the workload, the system provides cryptographic proof that only the approved code is running on genuine Confidential Computing hardware and that the environment has been properly hardened.

Real-World Medical AI Evaluation: MedPerf & GCP Confidential Computing Demo

Learn how ML Commons MedPerf integrates with Google Cloud Confidential Compute to enable secure, real-world evaluation of medical AI models.

From theory to critical impact: Advancing brain tumor research

This technology is already driving critical research through the Federated Tumor Segmentation (FeTS) initiative. Brain tumors, such as glioblastomas, are rare, making it difficult for any single hospital to collect enough data for high-accuracy AI training.

Compounding the problem, a model that performs perfectly in one hospital can struggle in another due to differences in patient demographics, data acquisition techniques, and even in equipment. 

Working with visionary researchers like Indiana University’s Dr. Spyridon Bakas, Northwestern University’s Dr. Yury Velichko, and the University of Alberta, Canada’s Dr. Amber Simpson, MedPerf on Google Cloud is validating AI models on private brain MRI data from around the world, and identifying potential performance gaps. 

For example, a model might be 95% accurate at one site but only 63% accurate at another. Our collaborative approach demonstrates that when an AI tool reaches a clinician, it has been proven to work across a truly representative patient population.

Achieving clinical trust and validation

The impact of this collaboration is best summarized by those on the front lines of clinical research. 

“My experience testing federated learning on Google Cloud has shown that the future of medical AI lies in secure, scalable, and collaborative cloud environments,” said Dr. Yury Velichko, associate professor, Radiology, Northwestern University. “Moving beyond the controlled lab setting to test these workflows in a production-ready infrastructure provided a unique opportunity to evaluate the performance and security of federated learning in real-world clinical applications.”

“Medical AI holds enormous promise for patients around the world, but that promise can only be realized if clinicians, researchers, and regulators can trust the benchmarks we use to evaluate it,” said Alexandros Karargyris, MedPerf lead, MLCommons. By bringing MedPerf onto Google Cloud’s Confidential Computing infrastructure, we have taken a major step toward a future where AI models can be rigorously tested on real patient data — without compromising privacy, intellectual property, or benchmark integrity.”

The future: Scaling secure medical breakthroughs

The collaboration between MLCommons and Google Cloud represents a fundamental shift toward privacy by design in healthcare AI. By making it easier to securely share and evaluate data and models, we are clearing the path for faster, safer, and more equitable medical breakthroughs. 

Research institutions and healthcare model developers interested in using the MedPerf platform on Google Cloud should contact medical@mlcommons.org or your Google Cloud account team.