Back to Production Traps
Data Substrates8 min read10 Oct 2026

The Small-File Lakehouse Explosion: Why Streaming Destroys Iceberg

Sub-minute streaming ingestion creates millions of uncompacted Parquet files in Apache Iceberg, destroying query speed. Learn how asynchronous bin-packing fixes it.

Author: Logic42 Sovereign Engineering Practice

Evaluating this architectural bottleneck in production?

The small-file lakehouse explosion is an operational crisis where high-frequency streaming pipelines write thousands of tiny Parquet files directly into Apache Iceberg or Delta tables. When data ingestion runs on 30-second commit cycles, table metadata trees swell to gigabytes, causing distributed query engines like Trino and DuckDB to time out during file discovery.

Data engineering teams across 2025 and 2026 pushed hard for real-time lakehouses. Business stakeholders demanded immediate data freshness for operational dashboards and agentic analytics. Instead of running hourly batch ETL jobs, teams connected Apache Flink or Kafka directly to cloud object storage. They committed data every 30 seconds.

In staging environments with small datasets, real-time streaming looks brilliant. You push an event, and it shows up in your SQL queries seconds later. But in production, continuous commits create an avalanche of tiny 120 KB Parquet fragments across dozens of partitions. Within two months, your clean table holds over 4,000,000 discrete objects.

Why do tiny Parquet files destroy lakehouse query performance?

Tiny Parquet files destroy query performance because distributed query engines spend 90% of their execution time reading object storage metadata instead of scanning column values. When an engine like Trino or ClickHouse initiates a query, it must parse thousands of Iceberg manifest files, resolve object URLs, and issue individual HTTP GET requests for every small file.

Look at how query latency and cloud storage costs escalate as uncompacted file counts rise:

File Size TierFiles per PartitionMetadata Snapshot SizeS3 LIST / GET RequestsTrino Planning LatencyEnd-to-End Query Time
512 MB (Target)4 files140 KB28 requests45 ms480 ms
128 MB (Acceptable)16 files520 KB112 requests120 ms1,250 ms
10 MB (Degraded)200 files6.8 MB1,400 requests850 ms4,600 ms
1 MB (Severely Fragmented)2,000 files64 MB14,000 requests4,200 ms18,900 ms
120 KB (Streaming Crisis)18,500 files4.8 GB140,000 requests38,000 msTIMEOUT (OOM)

Notice what happens at the cloud infrastructure layer. When your table reaches the streaming crisis tier, a single dashboard refresh triggers 140,000 S3 GET calls. AWS rate-limiting kicks in with HTTP 503 SlowDown errors. You pay $4,200 every month in redundant API fees while executive dashboards crash with OutOfMemory exceptions.

We saw this happen. A financial client ran continuous Flink ingestion into an Iceberg lakehouse. The engine gave up. The problem isn't the storage engine. It's the ingestion pattern.

The Three Hidden Costs of Lakehouse Fragmentation

Unmanaged streaming files don't just slow down ad-hoc queries. They undermine your entire data platform across three critical dimensions.

1. Object Storage Metadata Lockup

Cloud object stores like AWS S3 and Google Cloud Storage are not hierarchical POSIX file systems. When an engine lists 500,000 files in a partition, it must paginate through hundreds of metadata markers. Reading manifest files takes thirty times longer than actually evaluating the SQL filter.

2. Loss of Vectorized Columnar Pruning

Parquet achieves exceptional performance through dictionary encoding and columnar compression. In a 512 MB file, column chunks compress efficiently and min/max row group statistics allow the reader to skip entire data blocks. In a 120 KB file, headers and footers consume 40% of the byte count, completely eliminating compression benefits.

3. Manifest Tree Thrashing and Commit Conflicts

Every streaming micro-batch creates a new table snapshot. When multiple Flink workers attempt concurrent commits while an unoptimized compaction job runs, Iceberg's optimistic concurrency control triggers commit collisions. Batches fail, backpressuring your Kafka brokers.

Architectural Comparison: Direct Streaming vs. Bin-Packed Compaction

To maintain sub-second query performance without sacrificing real-time visibility, you must separate streaming ingest buffers from queryable analytical partitions.

The Small-File Lakehouse Explosion: Streaming vs. Bin-Packed Iceberg

The Logic42 Architectural Fix: Asynchronous Bin-Packing & Z-Order Sorting

Under our Build-Transfer-Operate practice, we deploy an asynchronous compaction worker pipeline that transparently consolidates streaming fragments without locking active readers.

# Logic42 Sovereign Lakehouse: Scheduled Iceberg Compaction Runner
from pyiceberg.catalog import load_catalog
from pyiceberg.table import Table

class SovereignIcebergOptimizer:
    def __init__(self, catalog_name: str, warehouse_uri: str):
        self.catalog = load_catalog(catalog_name, **{"warehouse": warehouse_uri})

    def execute_bin_packing(self, table_identifier: str, target_file_size_mb: int = 512):
        table: Table = self.catalog.load_table(table_identifier)
        target_bytes = target_file_size_mb * 1024 * 1024
        
        print(f"[OPTIMIZER] Analyzing snapshot lineage for {table_identifier}...")
        
        # Rewrite data files using bin-pack strategy with target byte boundaries
        rewrite_result = table.rewrite_data_files(
            strategy="binpack",
            options={
                "target-file-size-bytes": str(target_bytes),
                "min-file-size-bytes": str(int(target_bytes * 0.75)),
                "max-file-size-bytes": str(int(target_bytes * 1.25)),
                "max-concurrent-file-group-rewrites": "8"
            }
        )
        
        # Expire obsolete snapshots older than 3 days to reclaim object storage
        table.expire_snapshots(older_than_ms=3 * 86400 * 1000)
        
        print(f"[OPTIMIZER] Compaction completed successfully. Consolidated {rewrite_result.rewritten_data_files_count} files into {rewrite_result.added_data_files_count} clean blocks.")

Three Rules for Sovereign Lakehouse Ingestion

  1. Decouple Ingest Buffers from Analytical Gold Tables: Route high-speed streaming events to an append-only raw buffer or temporary staging partition. Don't write 30-second commits directly to core production datasets.

  2. Automate Continuous Bin-Packing: Run an independent Spark or PyIceberg worker that continually coalesces files smaller than 64 MB into optimal 512 MB columnar chunks. This keeps Trino file planning under 50ms.

  3. Prune Snapshot Lineage Aggressively: Retain table snapshots only as long as your rollback SLA requires. Dropping orphan manifests older than 72 hours reclaims terabytes of ghost storage on private MinIO or AWS S3 buckets.

If you don't manage small-file explosion, your real-time lakehouse will collapse under its own metadata weight. We engineer deterministic storage substrates directly inside your enterprise perimeter so your queries run fast, regardless of ingestion scale.

Sovereign Practice Diagnostic
4 Pillars · 16 Calibrated Checkpoints

Engineering Sovereign Data Boundaries?

Whether you are navigating cross-border CLOUD Act liability, implementing confidential compute enclaves, or eliminating vector decay, our data practice designs hardened substrates with client-held cryptographic custody.

Unlocks:Boardroom PDF DossierExcel Working PapersLegal Playbook (.md)
Confidential & Zero Third-Party Telemetry · Encrypted Practice Intake
Share this note
SUBSCRIBE TO FIELD NOTES

New Field Notes in your inbox.

We publish when we have something worth saying — reference architectures, benchmark tests, and engineering analysis. No cadence, no spam.