This is the first of a two-part series exploring how Google Cloud is bringing the foundational values of a high-performance parallel filesystem–TB/s throughput, sub-ms latency at high client scale, and POSIX support–to a broader set of use cases and users.
Historically, due to the cost and special purpose nature of parallel filesystems, colder data had to be stored outside of the filesystem and AI developers have had to maintain separate, slower environments for writing code, compiling libraries, and managing repositories. This fragmentation increases the toil of manual data staging, dataset copying, and managing disjointed namespaces.
Google Cloud Managed Lustre is solving these problems through our 6 cents/GB*month Dynamic Tier and by optimizing Managed Lustre performance for a range of development tasks and workloads – making Managed Lustre a “One-Stop Shop” for high-performance AI and HPC workloads.
Lower Cost: More Lustre for Less with the Dynamic Tier
The Managed Lustre Dynamic Tier provides sub-ms latency for hot data, which allows you to store all of your data in a single namespace, and costs only 6 cents/GB*month.
-
Throughput, capacity scale and client scale: Throughput scales linearly with capacity up to 80 PB, while sub-ms latency for hot data remains stable as you scale to tens of thousands of clients.
-
Single-flat fee: Predictable pricing. No independent charges for disk media types, data movement within the namespace, or metadata IOPS.
-
Read Latencies: Sub-ms latencies for High-Performance Cache (SSD). The Capacity Pool (“HDD”) is built on Google Cloud Hyperdisk throughput, which has an average read latency of 10 to 30 ms.
Recommended workloads for Dynamic Tier
-
Multi-Epoch Training and/or Training with Optimized Fetch Sizes: Hot data is promoted to the High Performance Cache (SSD) after the first run. Larger data prefetch will allow you to take advantage of the Dynamic Tier cost structure and gain from low-latency SSD.
-
Write-Heavy Checkpointing: Bursty checkpoint writes land directly in the High Performance Cache. Older checkpoints are transparently demoted to the Capacity Pool (HDD).
-
Rapid Checkpoint Restore: New checkpoints are written to the High Performance Cache, enabling low-latency checkpoint restores.
-
Interactive Snappiness for Developers: Low-latency tasks like git cloning, compiling libraries, or running notebooks benefit from a local-disk feel (~300µs average read latencies) on the same shared workspace hosting large training sets.
Frictionless development: Lustre as a one-stop shop for developer’s workloads
In addition to Managed Lustre’s scalability for large AI and HPC workloads (checkpoint/restart/data-loading), it also meets the demands for interactive work, meaning developers can start on Managed Lustre and stay on Managed Lustre throughout the entire workload lifecycle:
Unified Foundation & Interactive Performance
Consolidates the AI and HPC lifecycle into a single namespace, providing a “local disk” feel for interactive work (Read more about the latency benefits of Managed Lustre experienced by Salesforce and others).
-
Latency: ~300µs average read latency—delivering up to 4x better responsiveness than alternative distributed file systems.
-
Accelerated Setup: Untar the Linux kernel in ~2 minutes (4.7x faster than alternative file solutions), run a 20-worker parallel git clone of Python in ~40 seconds, compile Python in ~200s.


High-Concurrency Broadcast & Cluster Startup
Managed Lustre maximizes GPU ROI by preventing storage bottlenecks during cluster initialization.
When thousands of worker nodes attempt to read the exact same file simultaneously (such as a shared model checkpoint, base weights, or container layer), traditional distributed file systems can choke on localized hotspotting, leaving high-cost GPU clusters idle for minutes.
-
Improves Aggregate Throughput for a large number of clients reading the same file: Demonstrates a 67% improvement over alternative file solutions.
-
Parallel Loading: Imports libraries like PyTorch across 4,000+ processes in under 60 seconds.

Run One-Stop Shop Workflows for Yourself
Here is the code for the tests we’ve run, so that you can perform your own testing.
Low latency for interactive access
We used fio to emulate small, low-concurrency reads and writes:
1 Storage system specs: 500 MBps per TiB tier of Managed Lustre, 108,000 GiB capacity. Zonal Filestore at 102,400 GiB capacity. Average throughput of 36.7 GB/s to 2,048 client VMs reading the same 40 GiB file.
- code_block
- <ListValue: [StructValue([('code', '# Read workloadrnfio –ioengine=libaio –filesize=100M –ramp_time=2s \rn –runtime=2m –time_based –numjobs=1 –direct=1 –verify=0 –randrepeat=0 \rn –group_reporting –directory=~/LUSTRE_MOUNT \rn –name=randread –blocksize=4k –iodepth=1 –readwrite=randread \rn –buffer_compress_percentage=50rnrn# Write workloadrnfio –ioengine=libaio –filesize=100M –ramp_time=2s \rn –runtime=2m –time_based –numjobs=1 –direct=1 –verify=0 –randrepeat=0 \rn –group_reporting –directory=~/LUSTRE_MOUNT \rn –name=randwrite –blocksize=4k –iodepth=1 –readwrite=randwrite \rn –buffer_compress_percentage=50'), ('language', ''), ('caption', )])]>
Accelerated setup
How to run Linux untar
- code_block
- <ListValue: [StructValue([('code', '# Download a kernel tarballrnwget -P /tmp https://cdn.kernel.org/pub/linux/kernel/v5.x/linux-5.18.9.tar.xzrnrn# Extract the archive to the Lustre mountrnmkdir ~/LUSTRE_MOUNT/kernelrntar -C ~/LUSTRE_MOUNT/kernel -xf /tmp/linux-5.18.9.tar.xz'), ('language', ''), ('caption', )])]>
In the above use case, you will want to take care to avoid the metadata performance tax that can come from running as root (Namely, tar issues chown and chmod calls to make extracted files’ owner+permissions match the ones recorded in the archive.). If you still wish to run as root (and have verified that this approach is compatible with your setup), you may specify `--no-same-owner --no-same-permissions` in order to ensure that extracted files maintain root as owner and have root’s default file permissions. In other words, it makes extraction as root behave like extraction as non-root (by ignoring the owner+permissions in the archive).
How to run Python gitclone
- code_block
- <ListValue: [StructValue([('code', 'git config –global checkout.workers 20rnmkdir ~/LUSTRE_MOUNT/pythonrngit clone https://github.com/python/cpython.git ~/LUSTRE_MOUNT/python'), ('language', ''), ('caption', )])]>
How to run Python compile
- code_block
- /dev/nullrnmake > /dev/nullrnpopd’), (‘language’, ”), (‘caption’, )])]>
High scale distribution
Aggregate throughput for distributing one large file to many nodes
Run the below on each client VM:
- code_block
- <ListValue: [StructValue([('code', '# Start fio in server modernfio –server'), ('language', ''), ('caption', )])]>
Run the below on a selected client VM:
- code_block
- <ListValue: [StructValue([('code', "# Create a 40 GiB filernfio –name=job1 \rn –ioengine=libaio \rn –direct=1 \rn –buffer_compress_percentage=50 \rn –blocksize=4m \rn –iodepth=32 \rn –filesize=40g \rn –readwrite=write \rn –filename ~/LUSTRE_MOUNT/40gb_testrnrn# Create an fio job file for the read workloadrncat < /tmp/read.fiorn[job1]rnfilename=${HOME}/LUSTRE_MOUNT/40gb_testrnrw=readrnbs=4mrnexitall_on_error=1rnEOFrnrn# Run the read workload using all client VMs in ~/hostfilernfio –client ~/hostfile /tmp/read.fio”), (‘language’, ”), (‘caption’, )])]>
Parallel loading of libraries across many processes
Run the below on a selected client VM:
- code_block
- <ListValue: [StructValue([('code', '# Install PyTorch in a virtual envrnpython3 -m venv ~/LUSTRE_MOUNT/envrnsource ~/LUSTRE_MOUNT/env/bin/activaternpip3 install –upgrade piprnpip3 install torch torchvision torchaudio rndeactivaternrn# Import PyTorch on all client VMs in ~/hostfile, 4 processes per hostrnmpirun –allow-run-as-root –oversubscribe –hostfile ~/hostfile -N 4 \rn bash -c 'source ~/LUSTRE_MOUNT/env/bin/activate && python3 -c "import torch"''), ('language', ''), ('caption', )])]>
Looking ahead and next steps
By eliminating the manual data staging tax and lowering entry costs with the Dynamic Tier, Google Cloud Managed Lustre is evolving from an elite, single-purpose engine into a highly versatile, unified storage fabric for the entire AI lifecycle.
In the second part of this series, we will focus on upcoming object integration features. Stay tuned!
Next steps
-
Run the benchmarks yourself (if you haven’t already): Deploy a Google Cloud Managed Lustre instance using the Google Cloud console and run tests provided above to benchmark your own workloads.
-
Explore the Dynamic Tier: Read the Google Cloud Managed Lustre Documentation to learn more.
-
Stay tuned for Part 2: In the next installment of this series, we will dive deep into upcoming object integration features and how they further simplify AI and HPC storage.
-
Get started with centralizing your development-to-training lifecycle on Google Cloud Managed Lustre!