Webinar-Lottie.svg

lakeFS Acquires DVC, Uniting Data Version Control Pioneers to Accelerate AI-Ready Data

webcros.svg
Webinar: The Adoption Playbook for AI-Ready Data

Announcing the lakeFS Summer 2026 Release: Governance for AI-Ready Data

John NoonanJohn Noonan
Last updated on July 28, 2026

Table of Contents

Watch how lakeFS works!

As AI moves into production, organizations are under growing pressure to govern the data behind it.

Models are trained on petabytes of multimodal data. AI agents are modifying enterprise systems on their own. And new regulations (like the EU AI Act) are raising expectations around accountability, reproducibility, and access control. Yet for many teams, proving what data was used, who changed it, or whether policy was enforced still means piecing together evidence from multiple systems.

lakeFS addresses these governance and compliance requirements at the infrastructure level which makes it an integral part of how data is managed from the beginning – rather than adding governance piece-by-piece on top of individual tools and processes.

The lakeFS Summer 2026 Release builds on this infrastructure-based architecture and introduces a new set of governance capabilities built specifically for AI-ready data. Together, they help organizations define trusted datasets, enforce access policies, isolate workloads, automate data lifecycle management, and maintain a complete history of every change.

Here’s what’s new

Datasets: A New Semantic Layer for Trusted Data Providing Ultimate Flexibility to Curate, Permission, and Consume AI-Ready Data

Datasets give teams a new way to define, version, and govern logical groupings of data with semantic meaning, without requiring consumers to know which repository, path, or table the data actually lives in.

A dataset is a named, curated view over existing lakeFS data, spanning objects, prefixes, Iceberg tables, and namespaces. Every update creates an immutable version. This way teams can pin pipelines, reproduce results, and prove exactly which data backed a given decision. Changes go through review with a full version log, and required governance attributes are enforced at write time, not after the fact. Governed metadata, including owner, classification, tier, and retention, makes datasets easy to find and manage. Access can be granted through the dataset itself, independent of the underlying repository’s permissions. Datasets are consumed with zero-copy reuse through familiar lakeFS, S3-compatible, and Iceberg clients. So teams adopt them without changing how they already work.

For AI and data teams operating under regulatory scrutiny, this makes governance automatic and enforced at write time. As opposed to reconstructed after the fact.

Multi-Tenancy: Fully Isolated Environments in a Single Deployment

Multi-Tenancy lets organizations run isolated “mini-lakeFS” environments inside a single installation, scoped to a business unit, customer, or security domain.

For enterprises managing multiple teams, regions, or regulatory boundaries under one data infrastructure, Multi-Tenancy is built for compliance-ready architectures at scale, without standing up separate deployments for every domain that needs isolation.

A New Auditing Solution: Every Operation in a Single Query

lakeFS now offers the ability to stream every operation into an Iceberg table. This gives compliance teams a straightforward way to answer, “who changed this data, and when?” What used to require piecing together logs across systems is now a single query against a consumable, structured record of every operation.

Attribute-Based Access Control (ABAC) Extends Permissions to Individual Objects

lakeFS has long supported attribute-based access control scoped at the repository level. The Summer Release extends that control to individual objects. This way teams can permission specific users to access or read specific objects based on their metadata, not just the repository they live in.

For teams managing mixed-sensitivity data in shared repositories, this makes it possible to enforce least-privilege access at the level regulators actually care about (the object). This prevents splitting data into separate repositories just to control who sees what.

Branch and Object Lifecycle Management for Governed Data Deletion and Retention

Together with garbage collection, the Summer Release gives teams a new set of capabilities to manage deletion. It’s built for PII handling and regulatory compliance.

  • Branch Lifecycle Management applies policy-based rules to automatically clean up stale branches once they’re no longer needed. The frees the references they hold and keeps environments tidy without manual cleanup.
  • Object Lifecycle Management goes a step further, automating data deletion at the prefix level based on retention rules teams define. As such, data is removed on a defined schedule rather than left to accumulate indefinitely.

Most compliance frameworks require that data is both protected while it’s in use and deleted once it’s no longer needed. Manual deletion is slow, inconsistent, and hard to prove during an audit.

With Branch Lifecycle Management, Object Lifecycle Management, and garbage collection working together, teams get a complete, policy-driven system for managing deletion end to end, reducing storage costs while giving compliance teams a defensible, automated answer to how retention limits are enforced.

Plus a New Catalog UI for Iceberg Tables

lakeFS is also extending its UI to support Iceberg tables, expanding the multimodal capabilities of the lakeFS UI. Alongside objects and files, users can now browse Iceberg tables in the same interface, making it easy to explore structured and unstructured data together. 

Want to Learn More?

Sign up for our release webinar for demos and more information on our latest features.

Resources

Datasets: https://docs.lakefs.io/datasets/

Auditing: https://docs.lakefs.io/admin/auditing/#iceberg-audit-log

ABAC: https://docs.lakefs.io/security/rbac/#attribute-based-access-control-for-objects

Frequently Asked Questions

The lakeFS Summer 2026 Release is a set of new governance capabilities for AI-ready data. It adds Datasets, an audit trail streamed to an Iceberg table, object-level attribute-based access control (ABAC), automated branch and object lifecycle management, multi-tenancy, and a new Catalog UI.

A dataset is a named, curated, versioned view over existing lakeFS data. One dataset can span objects, prefixes, Iceberg tables, and namespaces, so teams consume governed data without knowing which repository or path it lives in. Every update creates an immutable version, and access can be granted through the dataset itself, independent of the underlying repository’s permissions.

 

It builds governance into how data is managed instead of bolting it on afterward. Governance attributes are enforced at write time, every change is versioned and reviewable, and operations can be streamed into a queryable audit trail. Teams define trusted datasets, enforce access policies, and can prove what data backed a given result.

Teams can now grant access to individual objects based on their metadata, not just at the repository level. That makes least-privilege access enforceable on mixed-sensitivity data inside a shared repository, without splitting data into separate repositories to control who sees what.

lakeFS streams every operation into an Iceberg table, so compliance teams get a structured, queryable record of activity. “Who changed this data, and when?” A question that used to mean piecing together logs across systems becomes a single query.

Multi-tenancy runs isolated “mini-lakeFS” environments inside one installation, each scoped to a business unit, customer, or security domain. Enterprises get compliance-ready isolation across teams, regions, or regulatory boundaries without standing up a separate deployment for every domain.

Two policy-driven capabilities work alongside garbage collection. Branch Lifecycle Management cleans up stale branches once they’re no longer needed; Object Lifecycle Management deletes data at the prefix level on a schedule teams define, rather than letting it accumulate indefinitely.

The Catalog UI extends the lakeFS UI to support Iceberg tables, bringing structured data alongside the objects and files already managed in lakeFS. Teams can explore multimodal data through a single interface.

It supports the obligations regulations like the EU AI Act raise — accountability, reproducibility, record-keeping, and access control. The queryable audit trail, immutable dataset versioning, object-level access control, and automated retention give high-risk AI systems the traceability and controls they increasingly require. It’s infrastructure that supports compliance, not a compliance guarantee on its own.

Sign up for the release webinar for demos and details.

 

We use cookies to improve your experience and understand how our site is used.

Learn more in our Privacy Policy