4 minute read

Introduction

This post is a short follow-up on the previous post, on OpenSharing with E-Series, which reaches the following conclusion:

OpenSharing and SANtricity

One of the things I mentioned in that post is that a solution stack with Versity S3 Gateway sharing E-Series tables or filesystems has some less obvious advantages due to the fact that unlike with S3-only storage, here I can run POSIX programs on data before creating OpenSharing shares.

That doesn't mean S3-only storage is doomed, but rather that this ability provides ways to solve common problems in unconventional ways.

Example with Tables: Iceberg table compaction

This example by Dremio explains why Iceberg tables need "defrag" (not really, but everyone understands that term).

The problem:

Consider a table fed by Apache Flink streaming with 1-minute checkpoint intervals. Each checkpoint creates new data files, often only 1-10MB each. After 24 hours of continuous streaming, you might have 1,440 tiny files instead of a handful of properly sized ones.

The solution is "defrag":

  • Read table files from target period (1 minute, 1 hour, 1 day, etc.)
  • Merge those tiny table into fewer larger (128-256 MiB) table files

And just like that, now that query that took 15 seconds again takes 1 second.

Notice how Dremio refers to 1 MiB files as "tiny". Well, folks, this is tiny and yes, there are people who use S3 like that…

Maybe you have a situation like that - thousands of table files getting created every day. Compacting them on E-Series should be painless. Why? Because E-Series is fast.

Table compaction scenario

Of course, we have to take a look, using a server and E-Series from the 2010s.

  • Create 10,000 table files on XFS
  • Compact (check several scenarios)
    • Try 4, 8 and 16 concurrent merges
    • Try 128 and 256 MiB target table size

The create step stressed the server out. This would be spread over time and data would be generated by clients, so the performance isn't very relevant to us. At 9:51:11am we see IO wait time jumps, which indicates storage bottleneck and that may be due to the way Spark was flushing from memory to disk (I haven't analyzed that).

Create 10K Iceberg tables

This thing runs in continuous loop, creating, compacting, dropping a bunch of times, so it's not trivial to check what's going on at every second without analyzing Spark itself.

PySpark

The compaction tests ran well. No I/O stalling, CPU (32 cores) fully utilized, write performance was hitting 2.5 GiB on a single volume (and I know this isn't the maximum write performance per volume; and notice we're reading tiny files here). None of this consumed CPU for S3 and TLS (if it was configured on VGW, and S3 used).

PySpark Iceberg Table Compaction

The last two runs, compacting into 256 MiB table files using 8 and 16 concurrent streams, respectively:

========================================================================

Dropping and recreating the table for a fresh run...
Generating and writing 10000 files...
Data generation complete.

--- Running Compaction with target-file-size-bytes=268435456, max-concurrent-file-group-rewrites=8 ---
+--------------------------+----------------------+---------------------+-----------------------+
|rewritten_data_files_count|added_data_files_count|rewritten_bytes_count|failed_data_files_count|
+--------------------------+----------------------+---------------------+-----------------------+
|                     10000|                    14|           3576929868|                      0|
+--------------------------+----------------------+---------------------+-----------------------+

+-------------------------+---------------------+
|rewritten_manifests_count|added_manifests_count|
+-------------------------+---------------------+
|                        2|                    1|
+-------------------------+---------------------+

Time Taken for compaction: 22.15 seconds

Files count: 14, Manifests count: 1

========================================================================

Dropping and recreating the table for a fresh run...
Generating and writing 10000 files...
Data generation complete.

--- Running Compaction with target-file-size-bytes=268435456, max-concurrent-file-group-rewrites=16 ---
+--------------------------+----------------------+---------------------+-----------------------+
|rewritten_data_files_count|added_data_files_count|rewritten_bytes_count|failed_data_files_count|
+--------------------------+----------------------+---------------------+-----------------------+
|                     10000|                    14|           3576879665|                      0|
+--------------------------+----------------------+---------------------+-----------------------+

+-------------------------+---------------------+
|rewritten_manifests_count|added_manifests_count|
+-------------------------+---------------------+
|                        2|                    1|
+-------------------------+---------------------+

Time Taken for compaction: 21.41 seconds

Files count: 14, Manifests count: 1

========================================================================

10,000 table files (~350 MiB) were compacted to 14 x 256 MiB tables in around 20 seconds.

I should have used 100x more data to make this exploration more "scientific", but I didn't think the conclusion would have been different.

Volume compaction scenario

That isn't really a thing, but I name it so because OpenSharing also supports Volume assets. Like Table assets, Volume assets may need "defragging" as well. In fact, the S3 tests with truly tiny files were created for precisely that kind of scenario - thousands of JSON files coming into S3 system every second.

One shouldn't do that; such files should be merged into tables before hitting S3, but stuff happens and the challenge for OpenSharing Volumes is the same: compact or convert data before making it available through OpenSharing Volumes.

I haven't tried doing similar tests with OpenSharing Volumes yet because that is super-niche now. But it's coming.

The ability to use both S3 and (BeeGFS) filesystem notifications makes it even easier to trigger and orchestrate such jobs, so I look forward to first OpenSharing Volumes opportunities.

Conclusion

With S3-only storage, it can be challenging to run many compacting jobs routinely, but we can do it with ease with Versity S3 Gateway and E-Series.

We can perform I/O-intensive slicing and dicing locally, and update OpenShare schemas and table names.

At 2 GB/s, more than 7 TB can be compacted per hour.

Note that there's no scale out. If you need 20 TB per hour, what then?

We'd have to run OpenSharing/Versity in scale-out way on BeeGFS (RWX) PVC volumes, as explained here. Or, we could run Hadoop 3 on E-Series and get scale-out that way.

But the objective isn't to "replace" S3 appliances or Hadoop clusters, but to use the ability to run "local" jobs on OpenSharing servers if that can help us solve problems that are otherwise difficult to solve. No one should have 128 KiB objects, but many have such (or worse) workloads.

Segregating such "bad" workloads onto isolated systems isn't a bad idea. Once tables have been compacted and made more usable, we can read them in an optimal way and use them without any impact on S3 infrastructure.