{"data":{"slug":"cactus-compute-cactus","name":"cactus","tagline":"Low-latency AI engine for mobile devices & wearables","github_url":"https://github.com/cactus-compute/cactus","owner":"cactus-compute","repo":"cactus","owner_avatar_url":"https://avatars.githubusercontent.com/u/196640840?v=4","primary_language":"C++","stars":5909,"forks":493,"topics":["ai","android","arm","edge","edge-ai","framework","ios","llamacpp","llm","llm-inference","llms","mobile","mobile-inference","on-device-ai","quantiz","rag","smartphone","speech","transformer","whisper"],"archived":false,"github_pushed_at":"2026-08-24T16:05:20+00:00","maintenance_label":"Very active","stars_delta_30d":374,"url":"https://www.graphcanon.com/tools/cactus-compute-cactus","markdown_url":"https://www.graphcanon.com/tools/cactus-compute-cactus.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/cactus-compute-cactus","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=cactus-compute-cactus","description":"Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots. ","homepage_url":"https://cactuscompute.com","license":"Other","open_issues":94,"watchers":46,"ai_summary":"Cactus is an AI framework optimized for low-latency operations on mobile and wearable devices. It supports a variety of tasks from speech recognition to general AI inference, particularly geared towards edge computing scenarios where quick response times are critical.","readme_excerpt":"# Cactus\n\n<img src=\"assets/banner.jpg\" alt=\"Logo\" style=\"border-radius: 30px; width: 100%;\">\n\n[![Docs][docs-shield]][docs-url]\n[![Website][website-shield]][website-url]\n[![GitHub][github-shield]][github-url]\n[![HuggingFace][hf-shield]][hf-url]\n[![Reddit][reddit-shield]][reddit-url]\n[![Blog][blog-shield]][blog-url]\n\nA hybrid edge-cloud AI engine for mobile devices & wearables.\n\n```\n┌─────────────────┐\n│  Cactus Engine  │ ←── OpenAI-compatible APIs for text, speech, and vision.\n└─────────────────┘     \n         │\n┌─────────────────┐\n│  Cactus Graph   │ ←── Zero-copy computation graph\n└─────────────────┘     \n         │\n┌─────────────────┐\n│ Cactus Kernels  │ ←── CPU/GPU kernels for (Apple, Samsung, Pixel, etc.)\n└─────────────────┘     \n         │\n┌─────────────────┐\n│ Cactus Quants   │ ←── Custom rotation-based quantization technique\n└─────────────────┘  \n```\n\n## Quick Demo (Mac)\n\n- Step 1: `brew install cactus-compute/cactus/cactus`\n- Step 2: `cactus run`\n\n## Cactus Engine\n\n```cpp\n#include \"cactus_engine.h\"\n\ncactus_model_t model = cactus_init(\n    \"path/to/weight/folder\",\n    \"path to txt or dir of txts for auto-rag\",\n    false\n);\n\nconst char* messages = R\"([\n    {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n    {\"role\": \"user\", \"content\": \"My name is Henry Ndubuaku\"}\n])\";\n\nconst char* options = R\"({\n    \"max_tokens\": 50,\n    \"stop_sequences\": [\"<|im_end|>\"]\n})\";\n\nchar response[4096];\nint result = cactus_complete(\n    model,            // model handle\n    messages,         // JSON chat messages\n    response,         // response buffer\n    sizeof(response), // buffer size\n    options,          // generation options\n    nullptr,          // tools JSON\n    nullptr,          // streaming callback\n    nullptr,          // user data\n    nullptr,          // pcm audio buffer\n    0                 // pcm buffer size\n);\n```\nExample response from Gemma4-E2B\n```json\n{\n    \"success\": true,        // generation succeeded\n    \"error\": null,          // error details if failed\n    \"cloud_handoff\": false, // true if cloud model used\n    \"response\": \"Hi there!\",\n    \"function_calls\": [],   // parsed tool calls\n    \"segments\": [],         // transcription segments (empty for chat)\n    \"confidence\": 0.8193,   // model confidence\n    \"confidence_threshold\": 0.7, // resolved handoff threshold (model-dependent)\n    \"time_to_first_token_ms\": 45.23,\n    \"total_time_ms\": 163.67,\n    \"prefill_tps\": 1621.89,\n    \"decode_tps\": 168.42,\n    \"ram_usage_mb\": 245.67,\n    \"prefill_tokens\": 28,\n    \"decode_tokens\": 50,\n    \"total_tokens\": 78\n}\n```\n\n## Cactus Graph\n\n```cpp\n#include \"cactus_graph.h\"\n\nCactusGraph graph;\nauto a = graph.input({2, 3}, Precision::FP16);\nauto b = graph.input({3, 4}, Precision::INT8);\n\nauto x1 = graph.matmul(a, b, false);\nauto x2 = graph.transpose(x1);\nauto result = graph.matmul(b, x2, true);\n\nfloat a_data[6] = {1.1f, 2.3f, 3.4f, 4.2f, 5.7f, 6.8f};\nfloat b_data[12] = {1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12};\n\ngraph.set_input(a, a_data, Precision::FP16);\ngraph.set_input(b, b_data, Precision::INT8);\n\ngraph.execute();\nvoid* output_data = graph.get_output(result);\n\ngraph.hard_reset(); \n```\n\n## Inference Speed\n\n- LLM: Gemma-4-E2B-CQ4 (1k-context prefill / decode for 100 tokens)\n- VLM: Gemma-4-E2B-CQ4 (256px image encode time / decode)\n- Transcribe: Parakeet-TDT-0.6B-CQ4 (20s audio end-to-end transcribe time)\n- 1k-Context RAM: peak MB during the LLM benchmark\n- No speculative decode or MTP, pure decode\n\nCommand: `cactus benchmark` [optional `--ios` or `--android`]\n\n| Device | LLM | VLM | Transcribe | RAM |\n|--------|-----|-----|------------|---------------|\n| Mac M5 Max | 2964tps / 154tps | 0.09s / 168tps | 0.15s | 1348MB |\n| Mac M4 Pro | 1963tps / 101tps | 0.25s / 112tps | 0.21s | 1225MB |\n| Mac M3 Pro | 1294tps / 64tps | 0.40s / 72tps | 0.37s | 735MB |\n| iPad/Vision Pro M5 | 1336tps / 71tps | 0.25s / 80tps | 0.27s | 703MB |\n| iPhone 17 Pro | 729tps / 37tps | 0.5s / 39tps | 0.51s | 644MB |\n| iPhone 15 Pro | 517tps / 26tps | 1.","github_created_at":"2025-04-23T14:33:43+00:00","created_at":"2026-07-11T11:42:42.871327+00:00","updated_at":"2026-08-24T18:01:15.223185+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"},{"slug":"speech-audio","name":"Speech & Audio","url":"https://www.graphcanon.com/categories/speech-audio","markdown_url":"https://www.graphcanon.com/categories/speech-audio.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/speech-audio"}],"tags":[{"slug":"ai","name":"ai"},{"slug":"android","name":"android"},{"slug":"arm","name":"arm"},{"slug":"edge","name":"edge"},{"slug":"edge-ai","name":"edge-ai"},{"slug":"framework","name":"framework"},{"slug":"ios","name":"ios"},{"slug":"llamacpp","name":"llamacpp"}],"trust":{"provenance":{"is_fork":false,"github_id":971447302,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-24T18:01:14.392Z","maintenance":{"label":"Very active","score":96,"methodology":"github_public_v1","releases_90d":3,"days_since_push":0,"last_release_at":"2026-08-17T17:42:06Z","stars_delta_30d":374,"open_issues_delta_30d":12},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:42:44.049Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-24T18:01:14.882Z"},"languages":{"value":["c++"],"source":"github.language","observed_at":"2026-08-24T18:01:14.882Z"},"license_spdx":{"value":"Other","source":"github.license","observed_at":"2026-08-24T18:01:14.882Z"}},"decision_facts":{"hosting":null,"pricing":{"model":"unknown"},"requirements":{"min_ram_gb":null,"requires_docker":false},"constraints":{"min_ram_gb":null,"pricing_model":"unknown","requires_docker":false},"when_to_use":["- When you need fast response times on mobile or wearable devices for tasks like speech recognition and general inference.","- For edge computing scenarios where power efficiency and quick decision-making are essential."],"when_not_to_use":["- In situations that require high-complexity AI applications beyond general inference, such as detailed image segmentation or extensive natural language understanding tasks.","- When working with desktop or server environments, as Cactus is specifically optimized for mobile and wearable hardware constraints."],"source":"enrich:decision_facts","observed_at":"2026-07-12T06:36:57.624Z"},"constraint_facets":{"min_ram_gb":null,"pricing_model":"unknown","requires_docker":false},"decision_summary":[{"label":"Pricing","value":"unknown"},{"label":"Adopt for","value":"Cactus - Low-latency AI engine optimized for mobile and wearable devices."},{"label":"License detail","value":"Other"}]}}