As AI moves into production, organizations are under growing pressure to govern the data behind it.
Models are trained on petabytes of multimodal data. AI agents are modifying enterprise systems on their own. And new regulations (like the EU AI Act) are raising expectations around accountability, reproducibility, and access control. Yet for many teams, proving what data was used, who changed it, or whether policy was enforced still means piecing together evidence from multiple systems.
lakeFS addresses these governance and compliance requirements at the infrastructure level which makes it an integral part of how data is managed from the beginning – rather than adding governance piece-by-piece on top of individual tools and processes.
The lakeFS Summer 2026 Release builds on this infrastructure-based architecture and introduces a new set of governance capabilities built specifically for AI-ready data. Together, they help organizations define trusted datasets, enforce access policies, isolate workloads, automate data lifecycle management, and maintain a complete history of every change.
Here’s what’s new
Datasets: A New Semantic Layer for Trusted Data Providing Ultimate Flexibility to Curate, Permission, and Consume AI-Ready Data
Datasets give teams a new way to define, version, and govern logical groupings of data with semantic meaning, without requiring consumers to know which repository, path, or table the data actually lives in.
A dataset is a named, curated view over existing lakeFS data, spanning objects, prefixes, Iceberg tables, and namespaces. Every update creates an immutable version. This way teams can pin pipelines, reproduce results, and prove exactly which data backed a given decision. Changes go through review with a full version log, and required governance attributes are enforced at write time, not after the fact. Governed metadata, including owner, classification, tier, and retention, makes datasets easy to find and manage. Access can be granted through the dataset itself, independent of the underlying repository’s permissions. Datasets are consumed with zero-copy reuse through familiar lakeFS, S3-compatible, and Iceberg clients. So teams adopt them without changing how they already work.
For AI and data teams operating under regulatory scrutiny, this makes governance automatic and enforced at write time. As opposed to reconstructed after the fact.
Multi-Tenancy: Fully Isolated Environments in a Single Deployment
Multi-Tenancy lets organizations run isolated “mini-lakeFS” environments inside a single installation, scoped to a business unit, customer, or security domain.
For enterprises managing multiple teams, regions, or regulatory boundaries under one data infrastructure, Multi-Tenancy is built for compliance-ready architectures at scale, without standing up separate deployments for every domain that needs isolation.
A New Auditing Solution: Every Operation in a Single Query
lakeFS now offers the ability to stream every operation into an Iceberg table. This gives compliance teams a straightforward way to answer, “who changed this data, and when?” What used to require piecing together logs across systems is now a single query against a consumable, structured record of every operation.
Attribute-Based Access Control (ABAC) Extends Permissions to Individual Objects
lakeFS has long supported attribute-based access control scoped at the repository level. The Summer Release extends that control to individual objects. This way teams can permission specific users to access or read specific objects based on their metadata, not just the repository they live in.
For teams managing mixed-sensitivity data in shared repositories, this makes it possible to enforce least-privilege access at the level regulators actually care about (the object). This prevents splitting data into separate repositories just to control who sees what.
Branch and Object Lifecycle Management for Governed Data Deletion and Retention
Together with garbage collection, the Summer Release gives teams a new set of capabilities to manage deletion. It’s built for PII handling and regulatory compliance.
- Branch Lifecycle Management applies policy-based rules to automatically clean up stale branches once they’re no longer needed. The frees the references they hold and keeps environments tidy without manual cleanup.
- Object Lifecycle Management goes a step further, automating data deletion at the prefix level based on retention rules teams define. As such, data is removed on a defined schedule rather than left to accumulate indefinitely.
Most compliance frameworks require that data is both protected while it’s in use and deleted once it’s no longer needed. Manual deletion is slow, inconsistent, and hard to prove during an audit.
With Branch Lifecycle Management, Object Lifecycle Management, and garbage collection working together, teams get a complete, policy-driven system for managing deletion end to end, reducing storage costs while giving compliance teams a defensible, automated answer to how retention limits are enforced.
Plus a New Catalog UI for Iceberg Tables
lakeFS is also extending its UI to support Iceberg tables, expanding the multimodal capabilities of the lakeFS UI. Alongside objects and files, users can now browse Iceberg tables in the same interface, making it easy to explore structured and unstructured data together.
Want to Learn More?
Sign up for our release webinar for demos and more information on our latest features.
Resources
Datasets: https://docs.lakefs.io/datasets/
Auditing: https://docs.lakefs.io/admin/auditing/#iceberg-audit-log
ABAC: https://docs.lakefs.io/security/rbac/#attribute-based-access-control-for-objects



