{"data":{"slug":"eventual-inc-daft","name":"Daft","tagline":"High-performance data engine for AI and multimodal workloads in Rust.","github_url":"https://github.com/Eventual-Inc/Daft","owner":"Eventual-Inc","repo":"Daft","owner_avatar_url":"https://avatars.githubusercontent.com/u/98941975?v=4","primary_language":"Rust","stars":5725,"forks":544,"topics":["ai-engineering","ai-pipeline","arrow","artificial-intelligence","big-data","data-engineering","distributed","distributed-computing","distributed-systems","embeddings","etl","huggingface","iceberg","machine-learning","multimodal","parquet","python","ray","rust"],"archived":false,"github_pushed_at":"2026-08-21T01:27:44+00:00","maintenance_label":"Very active","stars_delta_30d":76,"url":"https://www.graphcanon.com/tools/eventual-inc-daft","markdown_url":"https://www.graphcanon.com/tools/eventual-inc-daft.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/eventual-inc-daft","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=eventual-inc-daft","description":"High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale","homepage_url":"https://daft.ai","license":"Apache-2.0","open_issues":371,"watchers":32,"ai_summary":"Provides capabilities to process images, audio, video, and structured data at scale designed with ai engineering, big-data, and distributed computing needs in mind.","readme_excerpt":"|Banner|\n\n|CI| |PyPI| |Latest Tag| |Coverage| |Slack|\n\n`Website <https://www.daft.ai>`_ • `Docs <https://docs.daft.ai>`_ • `Installation <https://docs.daft.ai/en/stable/install/>`_ • `Daft Quickstart <https://docs.daft.ai/en/stable/quickstart/>`_ • `Community and Support <https://github.com/Eventual-Inc/Daft/discussions>`_\n\nDaft: High-Performance Data Engine for AI and Multimodal Workloads\n==================================================================\n\n|TrendShift|\n\n`Daft <https://www.daft.ai>`_ is a high-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale.\n\n* **Native multimodal processing:** Process images, audio, video, and embeddings alongside structured data in a single framework\n* **Built-in AI operations:** Run LLM prompts, generate embeddings, and classify data at scale using OpenAI, Transformers, or custom models\n* **Python-native, Rust-powered:** Skip the JVM complexity with Python at its core and Rust under the hood for blazing performance\n* **Seamless scaling:** Start local, scale to distributed clusters on `Ray <https://docs.daft.ai/en/stable/distributed/ray/>`_, `Kubernetes <https://docs.daft.ai/en/stable/distributed/kubernetes/>`_\n* **Universal connectivity:** Access data anywhere (S3, GCS, Iceberg, Delta Lake, Hugging Face, Unity Catalog)\n* **Out-of-box reliability:** Intelligent memory management and sensible defaults eliminate configuration headaches\n\nGetting Started\n---------------\n\nInstallation\n^^^^^^^^^^^^\n\nInstall Daft with ``pip install daft``. Requires Python 3.10 or higher.\n\nFor more advanced installations (e.g. installing from source or with extra dependencies such as Ray and AWS utilities), please see our `Installation Guide <https://docs.daft.ai/en/stable/install/>`_\n\nQuickstart\n^^^^^^^^^^\n\nGet started in minutes with our `Quickstart <https://docs.daft.ai/en/stable/quickstart/>`_ - load a real-world e-commerce dataset, process product images, and run AI inference at scale.\n\n\nMore Resources\n^^^^^^^^^^^^^^\n\n* `Examples <https://docs.daft.ai/en/stable/examples/>`_ - see Daft in action with use cases across text, images, audio, and more\n* `User Guide <https://docs.daft.ai/en/stable/>`_ - take a deep-dive into each topic within Daft\n* `API Reference <https://docs.daft.ai/en/stable/api/>`_ - API reference for public classes/functions of Daft\n\nBenchmarks\n----------\n|Benchmark Image|\n\nTo see the full benchmarks, detailed setup, and logs, check out our `benchmarking page. <https://docs.daft.ai/en/stable/benchmarks>`_\n\nContributing\n------------\n\nWe ❤️ developers! To start contributing to Daft, please read `CONTRIBUTING.md <https://github.com/Eventual-Inc/Daft/blob/main/CONTRIBUTING.md>`_. This document describes the development lifecycle and toolchain for working on Daft. It also details how to add new functionality to the core engine and expose it through a Python API.\n\nHere's a list of `good first issues <https://github.com/Eventual-Inc/Daft/issues?q=is%3Aopen+is%3Aissue+label%3A%22good+first+issue%22>`_ to get yourself warmed up with Daft. Comment in the issue to pick it up, and feel free to ask any questions!\n\nTelemetry\n---------\n\nTo help improve Daft, we collect non-identifiable data via Scarf (https://scarf.sh).\n\nTo disable this behavior, set the environment variable ``DO_NOT_TRACK=true``.\n\nThe data that we collect is:\n\n1. **Non-identifiable:** No session IDs or user identifiers are collected\n2. **Metadata-only:** We do not collect any of our users’ proprietary code or data\n3. **For development only:** We do not buy or sell any user data\n\nPlease see our `documentation <https://docs.daft.ai/en/stable/telemetry/>`_ for more details.\n\n.. image:: https://static.scarf.sh/a.png?x-pxid=31f8d5ba-7e09-4d75-8895-5252bbf06cf6\n\nRelated Projects\n----------------\n\n+---------------------------------------------------+-----------------+---------------+-------------+-----------------+-----------------------------+-------------+\n| Engine","github_created_at":"2022-04-25T22:02:29+00:00","created_at":"2026-07-11T11:28:37.418592+00:00","updated_at":"2026-08-22T00:01:32.567656+00:00","categories":[{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"},{"slug":"model-training","name":"Model Training","url":"https://www.graphcanon.com/categories/model-training","markdown_url":"https://www.graphcanon.com/categories/model-training.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/model-training"}],"tags":[{"slug":"ai-engineering","name":"ai-engineering"},{"slug":"ai-pipeline","name":"ai-pipeline"},{"slug":"arrow","name":"arrow"},{"slug":"artificial-intelligence","name":"artificial-intelligence"},{"slug":"big-data","name":"big-data"},{"slug":"data-engineering","name":"data-engineering"},{"slug":"distributed-computing","name":"distributed-computing"},{"slug":"embeddings","name":"embeddings"}],"trust":{"provenance":{"is_fork":false,"github_id":485548415,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-22T00:01:31.783Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":10,"days_since_push":0,"last_release_at":"2026-08-14T22:18:55Z","stars_delta_30d":76,"open_issues_delta_30d":29},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:28:38.655Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-22T00:01:32.244Z"},"has_cli":{"value":true,"source":"pyproject.toml:[project.scripts]","observed_at":"2026-08-22T00:01:32.244Z"},"languages":{"value":["rust","python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-22T00:01:32.244Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-22T00:01:32.244Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["- When you require high performance and efficiency in a multilingual environment, particularly if projects are primarily developed in Rust","- If your project involves heavy multimedia data handling, including images, audio, video alongside structured data"],"when_not_to_use":["- Avoid using Daft for projects where Python dominates the tech stack or development ecosystem","- When performance requirements are lower and ease of use is prioritized over speed"],"source":"enrich:decision_facts","observed_at":"2026-07-17T07:18:27.979Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"Daft is a Rust-based high-performance data engine for AI and multimodal workloads that supports processing various types of structured and unstructured data at scale."}]}}