pachyderm logo

pachyderm

pachyderm/pachyderm

Data-Centric Pipelines and Data Versioning

GraphCanon updated 2w · GitHub synced 2w

6.3k stars577 forksLast push 1y Go Apache-2.0

Decision brief

Pachyderm offers a robust platform for managing data-centric pipelines and data versioning with advanced features suitable for analytics and big-data processing in distributed systems.

Good fit when

  • If you need granular data lineage tracking within your projects, as Pachyderm ensures every transformation is captured.
  • For applications requiring reproducibility across experiments; Pachyderm provides versioned datasets that can be reliably re-used and compared.

Avoid when

  • If your organization does not require data versioning or cannot benefit from reproducibility features, such as for simple projects with minimal data mutation.
  • For scenarios where Docker container management overhead is undesirable; Pachyderm relies heavily on containers and Kubernetes, which might complicate smaller-scale workflows.
Pricing:
unknown - The repository does not specify detailed pricing, but as an open-source tool under the Apache-2.0 license, it is freely available for use and modification.
Requirements:
Min -1 GB RAM; Pachyderm deployment requires a Kubernetes cluster when deployed in production-scale environments.

Observed Jul 17, 2026 · Source: enrich:decision_facts

Verify the decision

Maintenance and security

Full trust report
Maintenance
Dormant (545d since push)
As of 2w
Provenance
Not a fork · Organization account
As of 2w
Security (OSV)
No lockfile
As of 1mo

Public GitHub metadata and optional OSV scans. Signals, not a guarantee. Trust methodology.

Install

go get github.com/pachyderm/pachyderm
pkg.go.dev

Similar tools

Same-category neighbours. No typed graph edges are catalogued for this tool yet.

Evidence and technical details

Sourced facts, taxonomy, compatibility claims, README excerpt, and machine-readable endpoints.

Overview

Pachyderm provides data-centric pipelines and manages data versioning within distributed systems, aiding in analytics and big-data processing.

Capability facts

Languages
go

Source: github.language · Aug 3, 2026

Categories

Tags

README

Getting Started

To start deploying your end-to-end version-controlled data pipelines, run Pachyderm locally or you can also deploy on AWS/GCE/Azure in about 5 minutes.

You can also refer to our complete documentation to see tutorials, check out example projects, and learn about advanced features of Pachyderm.

If you'd like to see some examples and learn about core use cases for Pachyderm:

For agents

This page has a .md twin and JSON over the API.

Was this helpful?

Anonymous feedback helps us improve pages and translations.