Webinar-Lottie.svg

lakeFS Acquires DVC, Uniting Data Version Control Pioneers to Accelerate AI-Ready Data

webcros.svg

Webinar October 22: Governing Autonomous AI Agents with lakeFS + MinIO

AI Data Platform: What It Actually Takes to Make Enterprise AI Data Reliable, Reproducible, and Audit-Ready

Tal SoferTal Sofer
Published October 8, 2026

•

Updated October 8, 2026

Table of Contents

Watch how lakeFS works!

Perhaps you’ve experience something like this before? It’s the week before a model risk review. A credit model has been in production since March, and the validator asks for one thing: the exact data it was trained on. The training config points at s3://ml-data/features/. Since March, three backfills have rewritten that prefix, a labeling fix has touched a third of the rows, and the engineer who ran the job has moved teams. What’s left is a notebook called features_final_v3.ipynb and a best guess. The model may be fine, but nobody can prove it, and the review stalls. That moment is the gap an AI data platform exists to close.

Organizations depend on enterprise data for AI but can’t say what that data looked like when a model or agent used it, and the gap is common. According to Dun & Bradstreet’s AI Momentum Survey of 10,000 businesses, 97% of organizations report active AI initiatives, but only 5% say their data is adequately ready to support them.

This article defines an AI data platform, covers where enterprise AI projects break without one, and shows what to put in place first when you plan your AI-data infrastructure.

What Is an AI Data Platform?

An AI data platform is the data infrastructure layer that gives AI models, agents, and data scientists governed, versioned, and reproducible access to enterprise data where it already lives. It records which version of the data each training run or agent action used, keeps changes isolated until they pass validation, and keeps an audit trail of who changed what.

The term gets stretched to cover GPU storage appliances and database suites. Here it means the data layer: what an AI workload sees, whether it’s correct, and whether you can prove both later. AI workloads need that layer because they read unstructured data (images, video, audio, PDFs) that doesn’t fit a table, because data scientists retrain on slightly different slices dozens of times, and because pipelines and agents write back to the same storage they read.

Data Warehouse
Data Lakehouse
AI Data Platform

Primary consumer

Analysts, BI dashboards

Analysts, data engineers, SQL engines

AI models, agents, data scientists, pipelines

Data it manages

Structured tables

Open table formats (Iceberg, Delta, Hudi) on object storage

Structured, semi-structured, and unstructured data together

Unit of history

Table time travel, retention-limited

Table snapshots, per table

Commits across files, tables, and metadata

“Reproducible” means

Re-running a query

Querying a table as of a snapshot

Re-creating the exact inputs of any training run or agent action

Who writes

Scheduled ETL

Pipelines

Pipelines, people, and autonomous agents

Most teams assemble an AI data platform from object storage they already run, a catalog, an orchestrator such as Airflow, and a version control layer that ties data state to each run. That’s one more component to operate and one more concept for data scientists to learn, a cost teams tend to accept after the first model they can’t reproduce.

Why AI Data Platforms Are Replacing Traditional Data Stacks

Traditional data stacks were built for analysts running known queries against curated tables. AI workloads read raw data, iterate on it constantly, and write outputs back into the same storage they read from. Once data changes underneath a model or agent, a stack that only tracks the current state can’t explain the results.

In a BI stack, data pipelines land clean tables, dashboards show the latest numbers, and overwriting a partition during a backfill is normal. Nobody needs last Tuesday’s table. AI breaks that. Data preparation never finishes, a model trained in March still serves in June, and feature pipelines, labeling vendors, and agents all write to shared storage. Each overwrite erases history a model depended on. The usual workaround, copying datasets before important runs, works until storage bills double and nobody remembers which copy fed which model.

An AI data platform doesn’t replace the warehouse, lakehouse, or object store. It adds a durable record of data state, isolation for anyone changing data, and checks before changes reach production. Our guide to AI-ready data architecture shows how these pieces fit together.

What an AI Data Platform Actually Covers

An AI data platform covers seven capabilities: storage and access across clouds, data versioning, quality enforcement, governance and audit trails, isolated access, centralized access without copies, and multimodal data management. Together, they make AI data reproducible, correct before it’s used, and provable after the fact, without moving it out of existing storage.

Diagram of an AI data platform layer between object storage and data scientists, pipelines, and AI agents, showing seven capabilities.

These capabilities depend on each other: versioning without isolation still lets bad writes land, and quality checks without versioning leave no way back. Our overview of AI-ready data management covers how they interact.

Scalable Data Storage and Access Across Clouds and Object Stores

AI training data is large, mostly unstructured, and read in parallel, so it lives in object storage: AWS S3, Azure Blob, Google Cloud Storage, or an S3-compatible store on premises. An AI data platform keeps object storage as the system of record and gives every consumer one consistent way to address it, with the read throughput GPU jobs need.

Data Versioning and Branching for Reproducible AI Runs

Data versioning records a dataset’s full state as an immutable commit; branching creates an isolated line of work from any commit. The commit ID then goes into each run’s metadata:

import mlflow

with mlflow.start_run():
    mlflow.log_param("git_sha", git_sha)
    mlflow.log_param("data_commit", data_commit_id)  # the exact data state this model trained on
    train(model, data_path)

The cost is retention: keeping history means keeping old objects, so you need a retention policy and garbage collection.

Data Quality Enforcement Before Data Reaches Training or Inference

New raw data lands somewhere isolated, automated checks run against it (schema, null rates, row counts, label distribution, leakage across splits), and it’s published only if they pass. This write-audit-publish pattern turns a failed check into a blocked change instead of a ticket filed after a model has already trained on bad data.

Governance, Audit Trails, and Compliance by Design

Every change records who made it, or which agent did, plus when and why, and every training run can be tied to a specific version. The workflow itself produces the evidence auditors ask for, and that regulations such as the EU AI Act require for high-risk systems, so nobody has to reconstruct it from logs.

Isolated Data Access for Agents, Experiments, and Teams

Each workload, from a data preparation job to an agent, gets its own view of production-scale data, and its changes stay invisible until merged. Copying a 200 TB dataset per experiment provides isolation too, at multiplied storage cost. Zero-copy branches do it through metadata. The trade-off is process: isolated work has to merge, and conflicting changes need a rule.

Centralized Access to Distributed Data Without Copying or Moving It

Sensor data in one cloud region, labeled data on premises, and feature tables in another account can sit behind one namespace with consistent permissions. Data access becomes a policy decision made once, instead of a copy request per job, and regulated data stays inside its required boundary.

Multimodal and Unstructured Data Management at Scale

A perception dataset might be millions of JPEGs, a JSON label file, and a Parquet metadata index. Table formats version only the table. An AI data platform commits images, labels, and metadata together, so a fix to 4,000 bounding boxes lands atomically with the images it describes.

The Data Infrastructure Gap That Breaks Enterprise AI Projects

Enterprise AI projects usually stall on six data problems: irreproducible training runs, agents overwriting production data, undetected bad data, no proof of what data drove a decision, idle GPUs, and slow data access. Their shared root cause is that nothing records what the data looked like when it was used, and nothing controls how it changes.

Failure Mode
What it Looks Like
What it Costs

Training runs that can’t be reproduced because input data changed

A retrain on “the same data” gives different metrics because the features table was backfilled in between.

Regressions can’t be bisected, and model comparisons stop meaning anything.

Agents corrupting or overwriting production data with no rollback path

An agent deduplicating customer records decides two different customers who share a name and city are one person, and deletes one in production.

Restore from backup, if one exists, and redo everything written since.

Bad data reaching production or training undetected

An upstream source switches a timestamp column from UTC to local time. Nothing fails; the model gets quietly worse.

It surfaces weeks later as a business metric, and tracing it means checking every input.

Compliance teams unable to prove what data was used in which AI decision

A regulator asks which records an underwriting model used to decline an applicant. The answer is a best guess.

Audit prep turns into months of team time.

Expensive GPUs sitting idle while waiting for data

Jobs copy data in from another region or re-fetch the same shards every epoch.

You pay for accelerators waiting on I/O. See our guide to GPU utilization.

Teams delayed by slow or restricted data access

A production slice for an experiment needs a ticket, a copy job, and a security review.

Fewer experiments run, and the copies escape governance.

Fixing these one at a time is how a data platform collects point solutions: backups for rollback, a lineage tool for audit, copy pipelines for access. A single layer that records data state and controls change addresses all six.

Common Pitfalls That Turn AI Data Platforms Into Technical Debt

AI data platforms become technical debt when teams put off five things: data versioning, isolation for agents and experiments, built-in compliance evidence, version-aware caching, and centralized access control. Each shortcut is cheap in a pilot and expensive in production, because retrofitting it means changing every pipeline, permission, and workflow already built on top.

Treating Data Versioning as an Afterthought Instead of Infrastructure

By the time things “stabilize,” dozens of data pipelines write to shared paths and notebooks hardcode bucket prefixes. Adding versioning then means touching all of them, and every model trained so far has no recorded inputs. Start with the first dataset, so a commit ID is part of every run from day one.

Letting Agents and Experiments Run Directly on Production Data

It’s the fastest way to get something working, and it’s how most first incidents happen. Staging copies go stale and cost storage. Zero-copy branches give agents and experiments real production data without write access to production, at the cost of a merge step someone or something has to approve.

Bolting Compliance Evidence On After the Fact Instead of Building It In

Evidence assembled on request means reconstructing history from logs never designed to answer the question. Built into the workflow, every change is a commit with an author and a message. The UK Home Office, the British government department, built data infrastructure that automates data quality, schema validation, and governance. Our overview of data compliance covers the controls auditors ask for.

Absorbing Performance Overhead From Uncached, Repeatedly-Fetched Data

Training reads the same data every epoch, sweep, and retrain. Uncached, that shows up as GPU idle time and egress charges. Caching fixes it only if the cache knows which version it holds; otherwise it feeds training stale data. Data preparation and training should read through a layer that caches by immutable version.

Managing Data Access Controls Separately Across Systems Instead of Enforcing Them Centrally

Bucket IAM policies, warehouse grants, filesystem ACLs, and a spreadsheet of who has access, last updated by someone who has since left, all drift independently. Agents make it worse, since each tends to run under a broadly permissioned service account. Central, role-based policies give one place to grant, review, and revoke access for people and agents.

How AI Data Platforms Are Being Used Across Industries

Industries adopt AI data platforms for the same core capabilities but for different reasons. Defense and life sciences need provable lineage for regulators. Automotive and manufacturing need reproducibility across huge, constantly changing unstructured data. Financial services needs isolated, governed data access for AI models that touch sensitive customer data.

Industry
Why reproducibility is non-negotiable
What the AI data platform provides

Defense and Government: Reproducibility and Audit Trails for Regulated AI

Classified or need-to-know data, air-gapped environments, and decisions that must be explainable down to the training data.

Lockheed Martin built an “AI factory” where teams share large sensitive datasets in a secure air-gapped environment, with reproducible experiments, access control, and complete lineage tracking.

GxP (good practice) data integrity rules, including the FDA’s 21 CFR Part 11 requirements for time-stamped audit trails on electronic records.

Immutable versions of every validation dataset, a full change history, and isolated branches so experiments never touch the validated state.

Automotive and Robotics: Version-Controlled Sensor Data at Scale

Camera, LiDAR, and telemetry datasets change constantly as teams ingest new drives and correct labels.

Volvo’s ML platform runs experiments on massive sensor and image datasets without data duplication, keeping experiments reproducible.

Manufacturing and Engineering: Version-Controlled Data for Model Development at Scale

Visual inspection data drifts with new suppliers, parts, and camera setups.

Datasets versioned alongside model code, so a spike in false rejects traces back to the data change behind it.

Financial Services: Isolated Data Access and Governance for Compliance-Sensitive AI

Model risk guidance such as the Federal Reserve’s SR 11-7 expects documented, validated models, including their data.

Isolated branches with role-based access for development, and a commit ID per training run for validators.

How AI Agents Change What Data Platforms Need to Do

AI agents change data platforms by turning AI from a reader of data into a writer of it. An agent that queries tables, generates code, and modifies datasets on its own needs more than access controls: its changes have to be isolated before they land, validated before they’re accepted, reversible when wrong, and tied to the exact data state each run saw.

Two runs of the same agent on the same task can take different steps and write different results, so the controls built for predictable pipelines don’t stretch to cover them. Our article on data agents covers how teams deploy them against enterprise data.

Agents Acting on Production Data Create Risk That Static Controls Cannot Catch

Access controls decide whether an identity may touch data. They can’t decide whether a change was correct. An agent with legitimate write access can still delete the wrong records or “fix” a column that wasn’t broken, and every write passes the permission check. Agents also write faster than any reviewer can read, so the answer has to be structural.

Agent Actions Need to Be Isolated, Versioned, and Reversible

Each agent task gets its own branch. Checks run when it finishes: pass, and the branch merges; fail, and it’s discarded without production ever seeing it. If a merged change proves wrong later, reverting to the previous commit is one operation. The trade-off is wiring agents to write to branches instead of direct paths.

Flow of an AI agent task on a data branch: checks pass and the branch merges, or checks fail and the branch is discarded.

Every Agent Run Needs a Reproducible Data State for Debugging and Compliance

The model version and prompt don’t reproduce an agent’s behavior if the data it read has changed. Record the commit it started from, the commit it produced, and the tool calls in between. Then you can rerun it after a fix, and show an auditor exactly what the agent knew when it acted.

lakeFS: The Control Plane for AI-Ready Data

lakeFS is the control plane for AI-ready data. It sits between object storage and the tools, agents, and users that consume that data, powered by a highly scalable data version control architecture. Data stays in place under your control, while every change becomes isolated, versioned, validated, and reversible.

Object storage stores objects reliably, but it doesn’t know which objects trained last month’s model, can’t keep an agent’s writes away from production, and can’t tell an auditor who changed a file and why. lakeFS adds that layer without replacing your warehouse, lakehouse, Spark, Airflow, or MLflow. Our guide on how to build infrastructure for AI-ready data shows where it fits.

Powered by Data Version Control: Branchable, Versioned, Rollback-Ready Workloads

Every experiment, training run, or agent action is tied to an exact, immutable version of the data, with versioning across data, code, and models, so any past result can be re-created, debugged, or built on with the same inputs. The workflow uses Git’s vocabulary:

# Create an isolated branch for the task
lakectl branch create lakefs://training-data/agent-task-142 --s lakefs://training-data/main

# Commit the result after the agent or job writes to the branch
lakectl commit lakefs://training-data/agent-task-142 -m "Deduplicate customer records (agent run 142)"

# Merge into main after checks pass
lakectl merge lakefs://training-data/agent-task-142 lakefs://training-data/main

If a merged change is wrong, an atomic rollback restores the previous state in seconds.

Isolation via Zero-Copy Branches for Agents, Experiments, and Teams

Branches write metadata, not data, so branching a petabyte-scale repository is instant. Pre-commit and pre-merge hooks run your validations on a branch, and a failing check blocks the merge. Branch protection rules keep anyone, or any agent, from writing to main directly.

Built-In Governance, Lineage, and Audit Trails

lakeFS captures data audit trails and lineage automatically across every workload, manages controlled and isolated data access across tools, agents, and users, reduces compliance risk with preventive data controls, and simplifies regulatory audits with built-in evidence instead of manual reporting. Role-based access control, SSO, and SCIM govern who can read, write, or merge where.

Centralized, Copy-Free Access Across Distributed and Multimodal Data

lakeFS gives tools, users, and AI agents centralized access to distributed data, so they can work with remote data as if it were local, across any back-end storage, with GPUs kept busy instead of waiting on data. It covers AWS S3, Azure Blob, Google Cloud Storage, S3-compatible, and POSIX storage, in cloud, on-premises, and air-gapped deployments, and manages images, logs, Parquet, Iceberg (via the REST catalog), Delta, and Hudi as one.

As David van Son, Software Engineer at Ellips, put it: “It used to take our entire ML engineering team 2 weeks to launch 2-3 new models. After implementing lakeFS, we now launch 6 new models in the same time with half the team.”

Conclusion

You don’t need to rebuild your stack to get an AI data platform working. Pick one training pipeline or agent workflow that touches production data and put it on versioned, branched data. Record the commit ID with every run, and route its writes through a branch with a validation check before merge.

That one workflow shows the team what changes: debugging starts from a known data state, a bad change comes back with a revert, and the next model review has its answer. The notebook named features_final_v3 can finally retire.

Then extend the pattern to the next pipeline, the next team, and your agents. lakeFS Enterprise runs it across petabyte-scale data, in your cloud, on premises, or air-gapped. Talk to the lakeFS team to see how it fits your environment.

Frequently Asked Questions

A data lakehouse combines data lake storage with warehouse-style table management, usually through open table formats like Iceberg or Delta. An AI data platform builds on that and adds what AI workloads need: versioning across structured and unstructured data, isolated environments for experiments and agents, validation before changes land, and an audit trail tying every model to its exact inputs. See our guide to AI-ready data architecture.

AI results depend on the data present when a job runs, and enterprise data changes constantly through backfills, corrections, and new ingestion. Without versioning, training runs can’t be reproduced, regressions can’t be traced to the data change behind them, and there’s no known-good state to roll back to. Learn more about data reproducibility.

They build compliance evidence into the workflow. The platform records every change with its author, time, and reason, ties every model to the exact data it used, and governs access through central policy. Validation before merge also stops non-compliant data from reaching production at all.

lakeFS sits between existing object storage and the tools, agents, and users that consume data, adding data version control, isolation, quality gates, and governance without moving data. It works across cloud, on-premises, and air-gapped deployments, and teams keep their current compute, orchestration, and ML tooling.

lakeFS records every change as an immutable commit, so each training run or agent action can be tied to the exact data state it read. Agents work on zero-copy branches, where teams validate, merge, or discard their changes without touching production. Read our practical guide to reproducible AI.