Webinar-Lottie.svg

lakeFS Acquires DVC, Uniting Data Version Control Pioneers to Accelerate AI-Ready Data

webcros.svg
lakeFS Named Cool Vendor™ in the 2026 Gartner® Coolest Vendor Innovations in Data Management

Data Engineering

Write-Audit-Publish for Data using lakeFS Hooks

lakeFS Hooks: Implementing Write-Audit-Publish for Data Using Pre-Merge Hooks

Write-Audit-Publish (continuous integration/continuous deployment of data) is the process of exposing data to consumers only after ensuring it adheres to best practices such as format,

PebbleDB SSTable

Concrete Graveler: Committing Data to Pebble SSTables

Introduction In our recent version of lakeFS, we switched to base metadata storage on immutable files stored on S3 and other common object stores.

Tiers in the Cloud

Tiers in the Cloud: How lakeFS caches immutable data on local-disk

Introduction We recently released the first version of lakeFS supported by Pebble’s sstable library – RocksDB. The release introduced a new data model which is

Ensuring Data Quality in a Data Lake Environment

Ensuring Data Quality in a Data Lake Environment

The quality of the data we introduce determines the overall reliability of our data lake. And the ingestion stage is a critical point for ensuring

AWS S3 Inventory

Getting Insights from Amazon S3 Inventory

Lists are everywhere. And there’s a good reason for that: lists have an order to them. They reduce our stress by assuring us that everything

Data Versioning as an Infrastructure

Why Data Versioning as an Infrastructure Matters

The demand for infrastructure that contributes to the collection, storage, and analysis of data is growing with the increasing amounts of data managed by organizations.

Monolith

Loosely Coupled Monolith vs Tightly Coupled Microservices

TL;DR With some thoughtful engineering, we can achieve a lot of the benefits that come with a microservice oriented architecture, while retaining the simplicity and

Data mesh

Data Mesh Applied: How to Move Beyond the Data Lake with lakeFS

The data mesh paradigm The Data Mesh paradigm was first introduced by Zhamak Dehghani in her article How to Move Beyond a Monolithic Data Lake

Object Storage: Everything You Need to Know

While Object Storage is not novel technology, it can still be overwhelming when getting started. Here’s a definitive guide to object-based storage with everything you

chaos data engineering

Chaos Data Engineering

Modern Data Lakes are a complexity tar pit. They involve many moving parts: distributed computation engines, running on virtualized servers connected by a software defined

System tests

System Tests: Lessons Learned From Developing For OSS Project

Overview In this article, I will try to cover some do’s and don’ts for system testing from the perspective of an open-source project. To keep

Building a data development environment

Building A Data Development Environment with lakeFS

Overview As part of our routine work with data we develop code, choose and upgrade compute infrastructure, and test new data. Usually, this requires running

[hubspot type=form portal=8040338 id=9f5646ec-3e20-4568-9d6e-b82fca022065]