{"data":{"slug":"vectifyai-pageindex","name":"PageIndex","tagline":"Document Index for Vectorless, Reasoning-based RAG","github_url":"https://github.com/VectifyAI/PageIndex","owner":"VectifyAI","repo":"PageIndex","owner_avatar_url":"https://avatars.githubusercontent.com/u/133959746?v=4","primary_language":"Python","stars":35204,"forks":3097,"topics":["agentic-ai","agents","ai","ai-agents","context-engineering","information-retrieval","llm","rag","reasoning","retrieval","retrieval-augmented-generation","vector-database"],"archived":false,"github_pushed_at":"2026-08-14T23:11:17+00:00","maintenance_label":"Very active","stars_delta_30d":1134,"url":"https://www.graphcanon.com/tools/vectifyai-pageindex","markdown_url":"https://www.graphcanon.com/tools/vectifyai-pageindex.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/vectifyai-pageindex","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=vectifyai-pageindex","description":"📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG","homepage_url":"https://pageindex.ai","license":"MIT","open_issues":157,"watchers":141,"ai_summary":"PageIndex is a document indexing system designed specifically for agentic AI tasks that require reasoning and retrieval-augmented generation (RAG) without relying on vector databases.","readme_excerpt":"### 🛠️ Deployment Options\n- **Self-host** — run locally with this open-source repo (using standard PDF parsing).\n- **Cloud Service** — production-grade pipeline with enhanced OCR, tree building, and retrieval for best results. Try instantly on our [Chat Platform](https://chat.pageindex.ai/), or integrate via [MCP](https://pageindex.ai/developer) or [API](https://pageindex.ai/developer).\n- **Enterprise** — dedicated or private deployment (VPC, on-prem). [Contact us](https://ii2abc2jejf.typeform.com/to/gVv7qkaN) or [book a demo](https://calendly.com/pageindex/meet) to learn more.\n\n---\n\n### 1. Install dependencies\n\n```bash\npip3 install --upgrade -r requirements.txt\n```\n\n---\n\n# Install optional dependency\npip3 install openai-agents","github_created_at":"2025-04-01T10:53:54+00:00","created_at":"2026-07-07T17:31:54.081296+00:00","updated_at":"2026-08-16T12:02:06.243536+00:00","categories":[{"slug":"ai-agents","name":"AI Agents","url":"https://www.graphcanon.com/categories/ai-agents","markdown_url":"https://www.graphcanon.com/categories/ai-agents.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/ai-agents"},{"slug":"data-retrieval","name":"Data & Retrieval","url":"https://www.graphcanon.com/categories/data-retrieval","markdown_url":"https://www.graphcanon.com/categories/data-retrieval.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/data-retrieval"}],"tags":[{"slug":"agentic-ai","name":"agentic-ai"},{"slug":"agents","name":"agents"},{"slug":"context-engineering","name":"context-engineering"},{"slug":"information-retrieval","name":"information-retrieval"},{"slug":"llm","name":"llm"},{"slug":"rag","name":"rag"},{"slug":"reasoning","name":"reasoning"},{"slug":"retrieval","name":"retrieval"}],"trust":{"provenance":{"is_fork":false,"github_id":958531089,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-16T12:02:05.459Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":5,"days_since_push":1,"last_release_at":"2026-08-14T18:25:13Z","stars_delta_30d":1134,"open_issues_delta_30d":18},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T10:58:01.778Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-16T12:02:05.918Z"},"languages":{"value":["python"],"source":"github.language+pyproject.toml","observed_at":"2026-08-16T12:02:05.918Z"},"license_spdx":{"value":"MIT","source":"github.license","observed_at":"2026-08-16T12:02:05.918Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["When your agentic AI project requires a non-vector database solution for reasoning and retrieval-augmented generation.","If you're working with tasks that need extensive context engineering and deep information retrieval without relying on traditional vector-based indexing."],"when_not_to_use":["For projects that strictly require the efficiency of vector databases, as PageIndex operates independently of these technologies.","When your application demands real-time indexing or quick data access methods that are more suited to vector database capabilities."],"source":"enrich:decision_facts","observed_at":"2026-07-11T12:56:23.771Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"PageIndex is a Python-based document indexing system that doesn't rely on vector databases. It's designed for agentic AI tasks where reasoning and context are key."}]}}