Tag: Apache Iceberg
All the articles with the tag "Apache Iceberg".
- 31 MIN READ•Jul 28, 2026
Why Iceberg V4 Wants to Retire Equality Deletes, and What Streaming Teams Should Do About It
Equality deletes made streaming upserts into Iceberg practical at the cost of read performance. V4 proposes retiring them in favor of deletion vectors with an async conversion path.
Apache IcebergStreamingData Engineering - 31 MIN READ•Jul 28, 2026
The Five Layers Between Your Lakehouse and a Trustworthy Agent
Agent reliability is a property of the stack the model sits on. Five layers with distinct owners and failure modes turn the agent is unreliable into a specific diagnosis.
AI AgentsApache IcebergData Architecture - 31 MIN READ•Jul 28, 2026
Apache Fluss and Kafka Solve Different Problems in an Iceberg Pipeline
Fluss puts a columnar, indexed hot tier between Kafka and Iceberg. Here's what it changes structurally, what Kafka still does better, and how to benchmark the comparison yourself.
Apache IcebergApache FlussKafka - 31 MIN READ•Jul 28, 2026
Serving Sub-Second Queries Over an Iceberg Lakehouse With a Hot Tier
A lakehouse cannot serve sub-second queries over seconds-old data. A hot tier in front solves it, with consequences for consistency, governance, and operational surface.
Apache IcebergStreamingData Serving - 31 MIN READ•Jul 28, 2026
Surviving Commit Conflicts When Dozens of Writers Hit the Same Iceberg Table
Commit conflicts multiply with writer count, and AI agents introduce unpredictable write patterns. Here's how to diagnose, tune, and architect around Iceberg's optimistic concurrency.
Apache IcebergConcurrencyData Engineering - 31 MIN READ•Jul 28, 2026
The Jackson 3 Problem in Apache Iceberg, and What It Means for Your Code
Jackson 3 changes everything: package names, unchecked exceptions, flipped defaults. Here's what breaks, why the engines are fine and your service isn't, and how to migrate safely.
Apache IcebergJacksonJava - 31 MIN READ•Jul 28, 2026
Wiring an AI Agent to Apache Polaris with the Model Context Protocol
The catalog is the right attachment point for AI agents working against a lakehouse. Here's how to wire the official Polaris MCP Server and add the read path it deliberately leaves out.
Apache IcebergMCPAI Agents - 31 MIN READ•Jul 28, 2026
Governing Iceberg Tables Across Regions Without Three Sets of Permissions
Catalog federation gives you one authorization model and one audit point across regions. Here's what it solves, what it doesn't, and how to build a topology you can actually govern.
Apache IcebergApache PolarisData Governance - 31 MIN READ•Jul 28, 2026
Federating Oracle With an Open Lakehouse Instead of Migrating It
Federate first so analytics work now, migrate what benefits from migrating, and leave the rest where it is indefinitely. Here's how pushdown and view layers make it work.
Apache IcebergOracleData Federation - 31 MIN READ•Jul 28, 2026
The Parquet Versioning Problem, and Why Iceberg Cares About It
Parquet files have a version field that doesn't reliably signal feature requirements. A new versioning discipline is coming, borrowing from Iceberg's format version model.
ParquetApache IcebergData Engineering - 31 MIN READ•Jul 28, 2026
Building Iceberg Pipelines in Python Without Standing Up Spark
A large share of production transformations fit comfortably on one machine. PyIceberg, DuckDB, and branch isolation give you a production path that debugs in an IDE.
Apache IcebergPythonPyIceberg - 31 MIN READ•Jul 25, 2026
Governing What Agents Cost You
Agents break the four assumptions analytics platforms were built on. A practical guide to identity, budgets, semantic layers, caching, and instrumentation for agent workloads.
AI agentscost governancedata platform