DataChain

DataChain is a Python-based AI-data warehouse for transforming and analyzing unstructured data like images, audio, videos, text and PDFs. It integrates with external storage (e.g. S3) to process data efficiently without data duplication and manages metadata in an internal database for easy and efficient querying.

Use Cases

ETL. Pythonic framework for describing and running unstructured data transformations and enrichments, applying models to data, including LLMs.
Analytics. DataChain dataset is a table that combines all the information about data objects in one place + it provides dataframe-like API and vecrorized engine to do analytics on these tables at scale.
Versioning. DataChain doesn't store, require moving or copying data (unlike DVC). Perfect use case is a bucket with thousands or millions of images, videos, audio, PDFs.

Getting Started

Visit Quick Start and Docs to get started with DataChain and learn more.

Key Features

📂 Multimodal Dataset Versioning.

Version unstructured data without moving or creating data copies, by supporting references to S3, GCP, Azure, and local file systems.
Multimodal data support: images, video, text, PDFs, JSONs, CSVs, parquet, etc.
Unite files and metadata together into persistent, versioned, columnar datasets.

🐍 Python-friendly.

Operate on Python objects and object fields: float scores, strings, matrixes, LLM response objects.
Run Python code in a high-scale, terabytes size datasets, with built-in parallelization and memory-efficient computing — no SQL or Spark required.

🧠 Data Enrichment and Processing.

Generate metadata using local AI models and LLM APIs.
Filter, join, and group datasets by metadata. Search by vector embeddings.
High-performance vectorized operations on Python objects: sum, count, avg, etc.
Pass datasets to Pytorch and Tensorflow, or export them back into storage.

Contributing

Contributions are very welcome. To learn more, see the Contributor Guide.

Community and Support

Docs
File an issue if you encounter any problems
Discord Chat
Email
Twitter

DataChain Studio Platform

DataChain Studio is a proprietary solution for teams that offers:

Centralized dataset registry to manage data, code and dependency dependencies in one place.
Data Lineage for data sources as well as derivative dataset.
UI for Multimodal Data like images, videos, and PDFs.
Scalable Compute to handle large datasets (100M+ files) and in-house AI model inference.
Access control including SSO and team based collaboration.

Name		Name	Last commit message	Last commit date
Latest commit History 411 Commits
.github		.github
docs		docs
examples		examples
src/datachain		src/datachain
tests		tests
.cruft.json		.cruft.json
.gitattributes		.gitattributes
.gitignore		.gitignore
.pre-commit-config.yaml		.pre-commit-config.yaml
CODE_OF_CONDUCT.rst		CODE_OF_CONDUCT.rst
LICENSE		LICENSE
README.rst		README.rst
mkdocs.yml		mkdocs.yml
noxfile.py		noxfile.py
pyproject.toml		pyproject.toml

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

DataChain

Use Cases

Getting Started

Key Features

Contributing

Community and Support

DataChain Studio Platform

About

Releases 61

Contributors 26

Languages

License

iterative/datachain

Folders and files

Latest commit

History

Repository files navigation

DataChain

Use Cases

Getting Started

Key Features

Contributing

Community and Support

DataChain Studio Platform

About

Topics

Resources

License

Code of conduct

Stars

Watchers

Forks

Releases 61

Contributors 26

Languages