{"data":{"slug":"ai-hypercomputer-jetstream","name":"JetStream","tagline":"Throughput and memory optimized engine for LLM inference on XLA devices","github_url":"https://github.com/AI-Hypercomputer/JetStream","owner":"AI-Hypercomputer","repo":"JetStream","owner_avatar_url":"https://avatars.githubusercontent.com/u/181000646?v=4","primary_language":"Python","stars":455,"forks":67,"topics":["gemma","gpt","gpu","inference","jax","large-language-models","llama","llama2","llm","llm-inference","llmops","mlops","model-serving","pytorch","tpu","transformer"],"archived":false,"github_pushed_at":"2026-01-05T17:51:16+00:00","maintenance_label":"Slowing","stars_delta_30d":4,"url":"https://www.graphcanon.com/tools/ai-hypercomputer-jetstream","markdown_url":"https://www.graphcanon.com/tools/ai-hypercomputer-jetstream.md","api_url":"https://www.graphcanon.com/api/graphcanon/tools/ai-hypercomputer-jetstream","graph_url":"https://www.graphcanon.com/api/graphcanon/graph?tool=ai-hypercomputer-jetstream","description":"JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).","homepage_url":null,"license":"Apache-2.0","open_issues":26,"watchers":22,"ai_summary":"JetStream is designed to optimize throughput and memory usage for the inference of large language models (LLMs) running on XLA-accelerated hardware like TPUs, with potential support for GPUs in future releases.","readme_excerpt":"> [!WARNING]\n> **Notice of Archival:** In an effort to streamline TPU inference efforts in open source, we have migrated core functionality in Jetstream to the new [tpu-inference](tpu.vllm.ai) repository. For this reason, we will be archiving Jetstream on February 1st 2026. Please note, archival does not mean deletion! Users will still be able to fork and clone Jetstream, we are simply shifting the repository to \"read-only\". To get Jetstream features and so much more, please check out [tpu.vllm.ai](tpu.vllm.ai).\n\n# JetStream is a throughput and memory optimized engine for LLM inference on XLA devices.\n\n## About\n\nJetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).\n\n## JetStream Engine Implementation \n\nCurrently, there are two reference engine implementations available -- one for Jax models and another for Pytorch models.\n\n### Jax\n\n- Git: https://github.com/google/maxtext\n- README: https://github.com/google/JetStream/blob/main/docs/online-inference-with-maxtext-engine.md\n\n### Pytorch\n\n- Git: https://github.com/google/jetstream-pytorch \n- README: https://github.com/google/jetstream-pytorch/blob/main/README.md \n\n## Documentation\n- [Online Inference with MaxText on v5e Cloud TPU VM](https://cloud.google.com/tpu/docs/tutorials/LLM/jetstream) [[README](https://github.com/google/JetStream/blob/main/docs/online-inference-with-maxtext-engine.md)]\n- [Online Inference with Pytorch on v5e Cloud TPU VM](https://cloud.google.com/tpu/docs/tutorials/LLM/jetstream-pytorch) [[README](https://github.com/google/jetstream-pytorch/tree/main?tab=readme-ov-file#jetstream-pytorch)]\n- [Serve Gemma using TPUs on GKE with JetStream](https://cloud.google.com/kubernetes-engine/docs/tutorials/serve-gemma-tpu-jetstream)\n- [Benchmark JetStream Server](https://github.com/google/JetStream/blob/main/benchmarks/README.md)\n- [Observability in JetStream Server](https://github.com/google/JetStream/blob/main/docs/observability-prometheus-metrics-in-jetstream-server.md)\n- [Profiling in JetStream Server](https://github.com/google/JetStream/blob/main/docs/profiling-with-jax-profiler-and-tensorboard.md)\n- [JetStream Standalone Local Setup](#jetstream-standalone-local-setup)\n\n\n# JetStream Standalone Local Setup\n\n## Getting Started\n\n### Setup\n```\nmake install-deps\n```\n\n### Run local server & Testing\n\nUse the following commands to run a server locally:\n```\n# Start a server\npython -m jetstream.core.implementations.mock.server\n\n# Test local mock server\npython -m jetstream.tools.requester\n\n# Load test local mock server\npython -m jetstream.tools.load_tester\n\n```\n\n### Test core modules\n```\n# Test JetStream core orchestrator\npython -m unittest -v jetstream.tests.core.test_orchestrator\n\n# Test JetStream core server library\npython -m unittest -v jetstream.tests.core.test_server\n\n# Test JetStream lora adapter tensorstore\npython -m unittest -v jetstream.tests.core.lora.test_adapter_tensorstore\n\n# Test mock JetStream engine implementation\npython -m unittest -v jetstream.tests.engine.test_mock_engine\n\n# Test mock JetStream token utils\npython -m unittest -v jetstream.tests.engine.test_token_utils\npython -m unittest -v jetstream.tests.engine.test_utils\n\n```","github_created_at":"2024-03-01T00:24:07+00:00","created_at":"2026-07-11T11:46:14.173791+00:00","updated_at":"2026-08-25T12:01:06.548457+00:00","categories":[{"slug":"inference-serving","name":"Inference & Serving","url":"https://www.graphcanon.com/categories/inference-serving","markdown_url":"https://www.graphcanon.com/categories/inference-serving.md","api_url":"https://www.graphcanon.com/api/graphcanon/categories/inference-serving"}],"tags":[{"slug":"gemma","name":"gemma"},{"slug":"gpt","name":"gpt"},{"slug":"gpu","name":"gpu"},{"slug":"inference","name":"inference"},{"slug":"jax","name":"jax"},{"slug":"large-language-models","name":"large language models"},{"slug":"llama","name":"llama"},{"slug":"llama2","name":"llama2"}],"trust":{"provenance":{"is_fork":false,"github_id":765455287,"owner_type":"Organization","methodology":"github_public_v1","parent_repo":null,"near_duplicate_slugs":[]},"computed_at":"2026-08-25T12:01:05.740Z","maintenance":{"label":"Slowing","score":36,"methodology":"github_public_v1","releases_90d":0,"days_since_push":231,"last_release_at":"2024-12-18T19:41:57Z","stars_delta_30d":4,"open_issues_delta_30d":1},"security_summary":{"status":"no_lockfile","scanner":null,"low_count":0,"high_count":0,"last_scan_at":"2026-07-11T11:46:15.327Z","medium_count":0,"scan_profile":"none","critical_count":0}},"capability_facts":{"scan":{"source":"repo_scan","observed_at":"2026-08-25T12:01:06.207Z"},"languages":{"value":["python"],"source":"github.language","observed_at":"2026-08-25T12:01:06.207Z"},"license_spdx":{"value":"Apache-2.0","source":"github.license","observed_at":"2026-08-25T12:01:06.207Z"}},"decision_facts":{"hosting":null,"pricing":null,"requirements":null,"constraints":null,"when_to_use":["* You are working with large language models (LLMs) that require efficient inference on hardware supported by XLA, particularly TPUs.","* Your project benefits from optimised memory usage and high throughput performance for inference tasks."],"when_not_to_use":["* If your primary compute platform is not an XLA-compatible device such as TPU; JetStream's current focus is on systems that are supported by XLA.","* When you need immediate support for GPUs, since GPU functionality is marked as a future potential enhancement."],"source":"enrich:decision_facts","observed_at":"2026-07-17T01:18:07.124Z"},"constraint_facets":null,"decision_summary":[{"label":"Adopt for","value":"JetStream optimises throughput and memory for LLM inference on XLA devices like TPUs, with potential GPU support in future."}]}}