The Small-File Lakehouse Explosion: Why Streaming Destroys Iceberg
Sub-minute streaming ingestion creates millions of uncompacted Parquet files in Apache Iceberg, destroying query speed. Learn how asynchronous bin-packing fixes it.
Author: Logic42 Sovereign Engineering Practice
The small-file lakehouse explosion is an operational crisis where high-frequency streaming pipelines write thousands of tiny Parquet files directly into Apache Iceberg or Delta tables. When data ingestion runs on 30-second commit cycles, table metadata trees swell to gigabytes, causing distributed query engines like Trino and DuckDB to time out during file discovery.
Data engineering teams across 2025 and 2026 pushed hard for real-time lakehouses. Business stakeholders demanded immediate data freshness for operational dashboards and agentic analytics. Instead of running hourly batch ETL jobs, teams connected Apache Flink or Kafka directly to cloud object storage. They committed data every 30 seconds.
In staging environments with small datasets, real-time streaming looks brilliant. You push an event, and it shows up in your SQL queries seconds later. But in production, continuous commits create an avalanche of tiny 120 KB Parquet fragments across dozens of partitions. Within two months, your clean table holds over 4,000,000 discrete objects.
Why do tiny Parquet files destroy lakehouse query performance?
Tiny Parquet files destroy query performance because distributed query engines spend 90% of their execution time reading object storage metadata instead of scanning column values. When an engine like Trino or ClickHouse initiates a query, it must parse thousands of Iceberg manifest files, resolve object URLs, and issue individual HTTP GET requests for every small file.
Look at how query latency and cloud storage costs escalate as uncompacted file counts rise:
| File Size Tier | Files per Partition | Metadata Snapshot Size | S3 LIST / GET Requests | Trino Planning Latency | End-to-End Query Time |
|---|---|---|---|---|---|
| 512 MB (Target) | 4 files | 140 KB | 28 requests | 45 ms | 480 ms |
| 128 MB (Acceptable) | 16 files | 520 KB | 112 requests | 120 ms | 1,250 ms |
| 10 MB (Degraded) | 200 files | 6.8 MB | 1,400 requests | 850 ms | 4,600 ms |
| 1 MB (Severely Fragmented) | 2,000 files | 64 MB | 14,000 requests | 4,200 ms | 18,900 ms |
| 120 KB (Streaming Crisis) | 18,500 files | 4.8 GB | 140,000 requests | 38,000 ms | TIMEOUT (OOM) |
Notice what happens at the cloud infrastructure layer. When your table reaches the streaming crisis tier, a single dashboard refresh triggers 140,000 S3 GET calls. AWS rate-limiting kicks in with HTTP 503 SlowDown errors. You pay $4,200 every month in redundant API fees while executive dashboards crash with OutOfMemory exceptions.
We saw this happen. A financial client ran continuous Flink ingestion into an Iceberg lakehouse. The engine gave up. The problem isn't the storage engine. It's the ingestion pattern.
The Three Hidden Costs of Lakehouse Fragmentation
Unmanaged streaming files don't just slow down ad-hoc queries. They undermine your entire data platform across three critical dimensions.
1. Object Storage Metadata Lockup
Cloud object stores like AWS S3 and Google Cloud Storage are not hierarchical POSIX file systems. When an engine lists 500,000 files in a partition, it must paginate through hundreds of metadata markers. Reading manifest files takes thirty times longer than actually evaluating the SQL filter.
2. Loss of Vectorized Columnar Pruning
Parquet achieves exceptional performance through dictionary encoding and columnar compression. In a 512 MB file, column chunks compress efficiently and min/max row group statistics allow the reader to skip entire data blocks. In a 120 KB file, headers and footers consume 40% of the byte count, completely eliminating compression benefits.
3. Manifest Tree Thrashing and Commit Conflicts
Every streaming micro-batch creates a new table snapshot. When multiple Flink workers attempt concurrent commits while an unoptimized compaction job runs, Iceberg's optimistic concurrency control triggers commit collisions. Batches fail, backpressuring your Kafka brokers.
Architectural Comparison: Direct Streaming vs. Bin-Packed Compaction
To maintain sub-second query performance without sacrificing real-time visibility, you must separate streaming ingest buffers from queryable analytical partitions.
The Logic42 Architectural Fix: Asynchronous Bin-Packing & Z-Order Sorting
Under our Build-Transfer-Operate practice, we deploy an asynchronous compaction worker pipeline that transparently consolidates streaming fragments without locking active readers.
# Logic42 Sovereign Lakehouse: Scheduled Iceberg Compaction Runner
from pyiceberg.catalog import load_catalog
from pyiceberg.table import Table
class SovereignIcebergOptimizer:
def __init__(self, catalog_name: str, warehouse_uri: str):
self.catalog = load_catalog(catalog_name, **{"warehouse": warehouse_uri})
def execute_bin_packing(self, table_identifier: str, target_file_size_mb: int = 512):
table: Table = self.catalog.load_table(table_identifier)
target_bytes = target_file_size_mb * 1024 * 1024
print(f"[OPTIMIZER] Analyzing snapshot lineage for {table_identifier}...")
# Rewrite data files using bin-pack strategy with target byte boundaries
rewrite_result = table.rewrite_data_files(
strategy="binpack",
options={
"target-file-size-bytes": str(target_bytes),
"min-file-size-bytes": str(int(target_bytes * 0.75)),
"max-file-size-bytes": str(int(target_bytes * 1.25)),
"max-concurrent-file-group-rewrites": "8"
}
)
# Expire obsolete snapshots older than 3 days to reclaim object storage
table.expire_snapshots(older_than_ms=3 * 86400 * 1000)
print(f"[OPTIMIZER] Compaction completed successfully. Consolidated {rewrite_result.rewritten_data_files_count} files into {rewrite_result.added_data_files_count} clean blocks.")
Three Rules for Sovereign Lakehouse Ingestion
-
Decouple Ingest Buffers from Analytical Gold Tables: Route high-speed streaming events to an append-only raw buffer or temporary staging partition. Don't write 30-second commits directly to core production datasets.
-
Automate Continuous Bin-Packing: Run an independent Spark or PyIceberg worker that continually coalesces files smaller than 64 MB into optimal 512 MB columnar chunks. This keeps Trino file planning under 50ms.
-
Prune Snapshot Lineage Aggressively: Retain table snapshots only as long as your rollback SLA requires. Dropping orphan manifests older than 72 hours reclaims terabytes of ghost storage on private MinIO or AWS S3 buckets.
If you don't manage small-file explosion, your real-time lakehouse will collapse under its own metadata weight. We engineer deterministic storage substrates directly inside your enterprise perimeter so your queries run fast, regardless of ingestion scale.
Engineering Sovereign Data Boundaries?
Whether you are navigating cross-border CLOUD Act liability, implementing confidential compute enclaves, or eliminating vector decay, our data practice designs hardened substrates with client-held cryptographic custody.
New Field Notes in your inbox.
We publish when we have something worth saying — reference architectures, benchmark tests, and engineering analysis. No cadence, no spam.