Home/Compare/ChainForge vs promptfoo

Comparison

ChainForge vs promptfoo

Verdict

Pick ChainForge if web-based visual programming tool for prompt testing on LLMs; supports local installation and Docker for advanced feature access; pick promptfoo if promptfoo aids in evaluating AI prompts, LLM agents, and RAG systems through declarative config testing with CI/CD support.

Markdown twin · ChainForge alternatives · promptfoo alternatives

GraphCanon updated today

ChainForge logo

ChainForge

ianarawjo/ChainForge

3.0kpushed Jun 10, 2026
vs
promptfoo logo

promptfoo

promptfoo/promptfoo

24kpushed Aug 1, 2026

Trust & integrity

SignalChainForgepromptfoo
Maintenance
Steady (70d since push)
As of today · github_public_v1
Very active (0d since push)
As of 2w · github_public_v1
Provenance
Not a fork · Personal account
As of today · github_public_v1
Not a fork · Organization account
As of 2w · github_public_v1
OSV dependency advisories
No lockfile (source not queried)
As of 1mo · osv@v1
No lockfile (source not queried)
As of 1mo · osv@v1
deps.dev advisories
Not queried
deps.dev@v1
Not queried
deps.dev@v1
OpenSSF Scorecard
Not queried
openssf-scorecard@v1
Not queried
openssf-scorecard@v1

Tagline

ChainForge
An open-source visual programming environment for battle-testing prompts to LLMs.
promptfoo
Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.

Stars

ChainForge
3.0k
promptfoo
24k

Forks

ChainForge
256
promptfoo
2.1k

Open issues

ChainForge
70
promptfoo
481

Language

ChainForge
TypeScript
promptfoo
TypeScript

Adopt for

ChainForge
Web-based visual programming tool for prompt testing on LLMs; supports local installation and Docker for advanced feature access.
promptfoo
promptfoo aids in evaluating AI prompts, LLM agents, and RAG systems through declarative config testing with CI/CD support.

Persona

ChainForge
-
promptfoo
-

Runtime

ChainForge
-
promptfoo
-

License

ChainForge
MIT
promptfoo
MIT

Last pushed

ChainForge
Jun 10, 2026
promptfoo
Aug 1, 2026

Categories

ChainForge
Evaluation & Observability, LLM Frameworks
promptfoo
Evaluation & Observability, LLM Frameworks

Trust and health

Maintenance

ChainForge
Steady (60%)
promptfoo
Very active (96%)

Days since push

ChainForge
70d
promptfoo
0d

Open issues (now)

ChainForge
70
promptfoo
481

Stars delta

ChainForge
+14 (30d)
promptfoo
Unknown

Open issues delta

ChainForge
+1 (30d)
promptfoo
Unknown

Owner type

ChainForge
User
promptfoo
Organization

Full report

ChainForge
Trust report
promptfoo
Trust report

Shared compatibility

  • Python · ChainForge: Python runtime · promptfoo: Python runtime

Choose ChainForge if…

  • Tags unique to ChainForge: ai, evaluation, large language models, llmops.
  • Need extensive features beyond web limitations
  • Leaner open-issue backlog (70).

When NOT to use ChainForge

  • Looking for a solution without local setup or Docker support
  • Prefer not managing API keys via environment variable configuration

Choose promptfoo if…

  • Tags unique to promptfoo: ci-cd, evaluation-framework, llm-evaluation, pentesting.
  • For comparing performance across GPT, Claude, Gemini, DeepSeek
  • More GitHub stars (24k vs 3.0k) - visibility, not fit.

When NOT to use promptfoo

  • If you do not require comparative analysis among multiple LLM models
  • If your project does not benefit from the specific red teaming capabilities offered by promptfoo

Explore

Sources

Every stat on this page traces to a dated GitHub sync, license file, enrichment field, or trust scan.

GitHub stars on cards: ChainForge 3.0k · promptfoo 24k (synced Aug 20, 2026).

Common questions

What is the difference between ChainForge and promptfoo?
ChainForge: An open-source visual programming environment for battle-testing prompts to LLMs.. promptfoo: Tool for evaluating prompts and AI agents by comparing performance across various models and red teaming.. See the comparison table for live GitHub stats and shared categories.
When should I choose ChainForge over promptfoo?
Choose ChainForge over promptfoo when Tags unique to ChainForge: ai, evaluation, large language models, llmops; Need extensive features beyond web limitations; Leaner open-issue backlog (70).
When should I choose promptfoo over ChainForge?
Choose promptfoo over ChainForge when Tags unique to promptfoo: ci-cd, evaluation-framework, llm-evaluation, pentesting; For comparing performance across GPT, Claude, Gemini, DeepSeek; More GitHub stars (24k vs 3.0k) - visibility, not fit.
When should I avoid ChainForge?
Looking for a solution without local setup or Docker support Prefer not managing API keys via environment variable configuration
When should I avoid promptfoo?
If you do not require comparative analysis among multiple LLM models If your project does not benefit from the specific red teaming capabilities offered by promptfoo
Is ChainForge or promptfoo more popular on GitHub?
promptfoo has more GitHub stars (23,838 vs 3,027). Stars measure visibility, not whether either tool fits your constraints.
Are ChainForge and promptfoo open source?
Yes - both are open-source projects on GitHub (ChainForge: MIT, promptfoo: MIT).
Where can I find alternatives to ChainForge or promptfoo?
GraphCanon lists graph-backed alternatives at ChainForge alternatives and promptfoo alternatives (ChainForge markdown twin, promptfoo markdown twin), ranked by typed relationship edges rather than popularity votes.
Is there a machine-readable version of this comparison?
Yes. The markdown twin at this comparison mirrors this page for agents and LLM crawlers, with the same stats table and FAQ answers.
Which is better maintained, ChainForge or promptfoo?
ChainForge: Steady. promptfoo: Very active. Compare maintenance labels, days since push, and release cadence in the trust section below - stars alone do not measure maintenance.
Where are the full trust reports for ChainForge and promptfoo?
GraphCanon publishes per-repo trust reports with dated maintenance, provenance, and scan summaries: ChainForge trust report; promptfoo trust report.

Was this helpful?

Anonymous feedback helps us improve pages and translations.