This post was originally published on this site

At Google Cloud, we know that you count on us to maintain the durability and integrity of your data at all times, both at rest and in transit. And now we’re making it easier for developers to take advantage of native data integrity features in Cloud Storage, by enabling end-to-end checksumming by default in all the Cloud Storage SDKs.

Like in any disk-based storage system, bits can flip anywhere in their journey, from the application all the way down to the disk. Cloud Storage has always let clients provide a checksum of the object data being uploaded, and receive a checksum of the data being downloaded. Also since its inception, Cloud Storage stores a checksum for every object in its metadata, regardless of how the object was uploaded into Cloud Storage.

But until recently, ensuring end-to-end data integrity required extra work on the part of developers to calculate and provide checksums to Cloud Storage. Cloud Storage always calculates the crc32 (32-bit cyclic redundancy check) of data it receives and ensures data stored on disk matches this checksum. When a client request includes the object’s checksum, Cloud Storage ensures that this checksum also matches. However, when an upload request doesn’t include a checksum, that upload is vulnerable to a bit flip while the data is in-flight, prior to the server-side checksum computation. Not all customers and clients enable client-side checksums by default, leaving data in this phase unprotected. 

To address this gap, the latest version of all Cloud Storage SDKs now internally checksums data being uploaded and passes this checksum to Cloud Storage, if it’s not provided by the application. The SDKs also support verifying the object’s checksum when an object is being downloaded.

Finally, there are many use-cases where applications download select ranges of objects instead of the full object. When using Cloud Storage SDKs with our gRPC API to perform a range read, the SDKs take advantage of gRPC’s built-in end-to-end range checksum, using it to verify the data it receives.

We highly recommend updating to our latest SDK versions to take advantage of these important integrity features.

And now, let’s peek under the hood

Ensuring continuous “chain-of-custody” between the data and its associated checksum from your application down to the disk platter, with no gap where a bit flip could go unnoticed, is quite challenging. And it’s critical to get this right: at our current scale of hundreds of thousands of Cloud Storage frontends, bit flips aren’t theoretical and do happen from time to time.

For instance, consider this simple example: when Cloud Storage receives your data in its frontend, this data gets encrypted with per-object encryption keys. This involves a data copy: the plaintext data is passed through an encryptor into a new memory buffer containing ciphertext. Extremely rarely, a bit in the source or destination memory buffer flips during this process. However, at our scale, extremely rare things happen routinely. 

In this situation, we maintain chain-of-custody by reversing the whole process: after encrypting the data (1), we calculate a checksum that protects the ciphertext. Then we decrypt the ciphertext (2), and if the resulting plaintext doesn’t match the original (3), we throw everything away and start over. This adds up to a lot of extra CPU time spent on encryption and checksumming, but it’s a necessary step to ensure data integrity.

1

Another challenge is how data gets broken up and aggregated as it passes through layers of our stack. As data gets uploaded to Cloud Storage, it gets split up into chunks, each of which has its own checksum. To manage data efficiently at scale, Cloud Storage groups thousands of chunks together into a storage unit we call a shard file. These gigabyte-sized files are how Cloud Storage ultimately delivers data to our cluster-level storage system, Colossus. Internally, Colossus uses Reed Solomon encodings to spread data across many disks and protect against the failures of individual disks, machines, and racks. This requires chopping up the shard file data into blocks, each of which is again protected by a checksum.

To maintain chain-of-custody of the data as it goes through all these transformations, we take advantage of some nifty properties of cyclic redundancy checks (CRCs), for example, concatenation. When you have two data buffers that each have their own CRC, you can cheaply compute the CRC of the two concatenated buffers without having to re-checksum the data. This comes in handy in many situations, such as when concatenating chunks together into shard files: Colossus can cheaply determine the CRC of the entire shard file from its constituent chunks and store that in its metadata.

Ultimately, the data lands on disks managed by our “D” file server (our network attached disks). D stores inline checksums for each range of data within a Colossus block. Whenever data is read from the disk, it is verified at several layers: The Colossus client verifies the data it reads against D’s inline checksums, and the Cloud Storage frontend reads data chunk-by-chunk, verifying each chunk against its checksum before sending it to the client. These chunk-level checksums are what enable our gRPC protocol to provide a checksum for a range read that can be verified by our SDKs, all without losing chain-of-custody.

2

Chain of Custody: Maintaining Data integrity across Data transformations

The above image shows the data integrity handoff across multiple layers under the hood of Google Cloud Storage. On reads, checksums are verified inline at multiple layers to prevent silent corruptions.

  1. Client passes full object checksum to Cloud Storage Frontends.

  2. Data is split into chunks and individual chunk level checksums are computed.

  3. Shard level checksums are computed based on concatenated chunk level CRCs.

  4. Shards are stored across disk blocks with another level of block level inline checksums.

Here on the Cloud Storage team, we remain dedicated to maintaining the highest standards of data integrity for our customers. By making end-to-end checksumming the default in our SDKs and maintaining chain-of-custody throughout our internal storage stack, data remains exactly as intended from the moment of upload to the final download. This continuous vigilance reflects our commitment to protecting your data at any scale. To take full advantage of these protections, we recommend updating to the latest version of our SDKs.