EvaliQA vs Promptfoo for evaluating AI systems

Promptfoo is a free open-source CLI for testing prompts and red teaming from a config file. We compare it with EvaliQA on test cases, judges, multi-turn runs, red teaming, production monitoring and reporting, and say where each one is the right pick.

A team that ships an AI system usually meets Promptfoo first. It is free, it runs from a terminal, and an engineer can have a prompt scored against 20 test cases in an afternoon without asking anyone for budget. The trouble starts later, when evaluation has to become something the whole team does rather than something one engineer runs. The config file lives in one repository, the results live on one laptop, the red teaming findings sit in a report nobody outside engineering opens, and nothing connects what the tests say to what users see in production. That gap between a tool one person runs and a practice a team owns is the problem this post is about.

We build EvaliQA, so read this as the opinion of one side. We have tried to keep it honest in two ways. Every Promptfoo fact below comes from Promptfoo's own documentation, repository or announcements, and each one is linked so you can check it. And we name the cases where Promptfoo is the right tool and EvaliQA is not. The short version is that Promptfoo is an excellent evaluation and red teaming CLI for a developer working inside a repository, and EvaliQA is a platform for a team that needs datasets, judges, multi-turn runs, red teaming, production traces and reports in one place, scored with the same metrics everywhere. The rest of the post goes through where that difference shows up and where it does not matter.

What the two tools are

Promptfoo describes itself as a tool to "test your prompts, agents, and RAGs" with red teaming, pentesting and vulnerability scanning for AI, and it is MIT licensed. You describe your prompts, providers and test cases in a declarative configuration file, run an evaluation from the command line, and read the results as a matrix view that puts every prompt and model combination side by side in a local web UI. It supports OpenAI, Anthropic, Azure, Google, Hugging Face, open models and custom API providers, and it is built to sit in CI. On 9 March 2026 Promptfoo announced it was joining OpenAI, with a commitment that the project stays open source and keeps supporting a range of providers. The same announcement cites more than 350,000 developers and 130,000 monthly active users, which is a real installed base and the reason most teams start here.

EvaliQA is a hosted evaluation platform organised around projects, datasets, test plans, eval runs and live traces. A project holds one AI system. A test plan binds a channel (HTTP or voice), one of 4 modes (single-turn evaluation, multi-turn evaluation, single-turn red teaming, multi-turn red teaming), a judge model, a set of metrics and one or more datasets. A run executes the plan and lands on a results page with per-metric pass rates, per-row judge reasoning, cost and latency. The same metric catalogue scores production traffic through the SDK or any OpenTelemetry exporter, and alerts fire when a live score slides. You bring your own model keys, and EvaliQA never bills for tokens on top.

The difference that decides most of what follows is that Promptfoo is a tool you run and EvaliQA is a system of record. A Promptfoo run produces a result in a local store, with a default retention of 30 days for traces, and sharing across a team is an Enterprise feature. An EvaliQA run is a row under its test plan that anyone in the workspace can open, compare and report on, and the dataset it ran against is versioned so the run stays reproducible. Neither design is wrong. They answer different questions, and the question you are asking should pick the tool.

Where Promptfoo is the better choice

Promptfoo wins when the unit of work is a prompt in a repository. Its matrix view runs several prompts against several models in one command and shows the grid, so comparing 3 prompt variants on 2 models is a single run. In EvaliQA each run of a plan has one target, so the same experiment is 6 runs and the Compare view on the plan page. If your daily work is prompt A/B testing with the config next to the code, Promptfoo is simply faster at that loop and we would not argue otherwise.

Promptfoo also has a wider set of deterministic checks and lets you write your own in code. Its assertion list includes javascript, python, ruby and webhook assertions alongside equals, contains, regex, is-json, latency and cost, and each assertion carries a weight. EvaliQA's deterministic metrics are Regex Match, JSON Schema, Length Check and Contains, with a few more in the catalogue, and anything that needs custom logic goes through an LLM judge rubric rather than a function you wrote. For a team with strong opinions encoded in code, that is a real constraint, and we say so.

A third point is the judge. Promptfoo lets each model-graded assertion name its own provider, so one assertion can be graded by a cheap model and another by a strong one. In EvaliQA the judge model is picked once per plan and applies to every LLM metric on it. We think the pin is the right default, because a judge is part of the measurement instrument and rotating it breaks run-to-run comparison, but it is less flexible. Finally, the Community edition runs locally or self-hosted at no cost, with one caveat covered under red teaming below. If data cannot leave your machines and you have no budget, Promptfoo is the answer and EvaliQA's self-hosted deployment is a Custom plan conversation.

Where test cases come from

In Promptfoo the test cases are part of the configuration. An engineer writes them, or loads them from a file, and they version with the code. That is tidy for a prompt, and it is also why evaluation tends to stall at the size one engineer can maintain by hand. Nobody outside the repository can add a case, a QA engineer cannot edit a row without a pull request, and a failure seen in production has no path into the test set except someone copying it over.

EvaliQA treats the dataset as its own artefact with 4 sources. You hand-write rows for a golden set. You generate rows from the project's description, business scenarios and out-of-scope lists, which is what Quick eval does in about 2 minutes from a pasted cURL command, producing between 1 and 50 cases scored on answer relevancy, bias, toxicity and task success. You import CSV or JSONL. And you promote a production trace into a dataset with Add trace to dataset, mapping the trace's input and output onto the row and typing the expected output you now know the agent should have produced. Datasets are versioned, so Duplicate as new version before a bulk change keeps old runs reproducible against the frozen rows.

The consequence is who can own the test bank. In EvaliQA a QA or prompt engineer curates rows in the editor, tags them with category, difficulty, persona and source, and the engineer never touches a YAML file. The dataset methodology we recommend, roughly 50 percent happy path, 20 percent edge cases, 10 percent metamorphic pairs, 10 percent forbidden topics and 10 percent inappropriate usage, is something you can see and rebalance on a page rather than infer from a file. What this does not solve is judgement. A generated row still needs a human to decide whether the expected answer is right, and the docs are blunt that you should review the first batch before trusting it.

Metrics and the judge model

Both tools split metrics into checks that need no model and checks that ask a model to grade. Promptfoo's model-assisted assertions include llm-rubric, g-eval, answer-relevance, context-faithfulness, context-recall, context-relevance, factuality and select-best, and each test can set a combined weighted threshold. EvaliQA's catalogue has 5 families, RAG and general LLM-judge metrics such as Answer Relevancy, Faithfulness and Contextual Precision and Recall, Agent metrics such as Task Success Rate and Tool Correctness, Security metrics such as PII Leakage and Prompt Injection Resistance, Deterministic metrics, and Custom metrics built with G-Eval or Custom Eval. Every metric produces a score from 0 to 1 and a threshold turns it into a pass or fail verdict, with a row passing only when every metric on the plan passes.

The practical difference is where the metric lives. A Promptfoo assertion belongs to a test case in a config. An EvaliQA metric belongs to the plan, so one dataset can be scored by two plans with different metric sets, and a custom metric written in plain language is saved as a workspace preset that any plan can attach. The same preset can then be bound to live traffic in online evaluation, which is the point we keep returning to. A number on a production dashboard means the same thing as the number on the release run because it is literally the same metric with the same threshold.

The pinned judge deserves one more paragraph because it is the choice people argue about. EvaliQA's docs are explicit that two identical runs will not produce identical numbers. On a 100-row plan with a strong judge, per-metric pass rates typically wobble by 2 to 3 percentage points between runs and aggregate pass rate by 1 to 2. That noise floor is the reason the judge is set once per plan rather than per metric. If each assertion could pick its own grader and someone changed one, the historical numbers would quietly stop being comparable and nobody would know which run introduced the shift. Promptfoo's per-assertion provider is more flexible, and it puts that discipline on you.

Multi-turn evaluation

Single-turn scoring is where both tools are closest, and multi-turn is where they diverge most. Promptfoo offers a simulated user provider inspired by the tau-bench benchmark. You give it instructions describing the user's goal and behaviour, set a maximum number of turns with a default of 10, optionally seed it with initial messages, and it talks to your agent until the turn limit, a stop token, or an error. It assumes the target accepts messages in the OpenAI chat format with role and content fields. It is a capable primitive and a reasonable way to get a conversation scored.

EvaliQA's multi-turn mode has 3 strategies because reproducibility and realism pull in opposite directions. Simulation replays pre-written scenario seeds, so the same seed against the same target gives the same conversation and a run-to-run comparison means something. Scripted replays an exact sequence of user turns, which is how you reproduce a production incident verbatim. Adaptive lets the Platform AI agent improvise the user side in response to what your agent said, which is the realistic one and also the noisy, expensive one. A row is one conversation with a scenario, a persona, an initial state, an expected outcome and a maximum turn count, and the generator draws on 15 persona types including Impatient, Confused, Non-native speaker and Manipulative.

The scoring side matters as much as the simulation. Conversations are scored by Agent metrics that read the whole exchange rather than one message, Task Success Rate, Goal Achievement, Role Adherence, Knowledge Retention, Conversational Flow, Repetitive Pattern Detection and Tools Error. The failure pattern the docs call out is worth knowing. A system that passes in Simulation and fails in Adaptive handles a script but cannot handle a user who deviates, and no single-turn metric would ever show you that. The honest limit is cost. A 40-conversation run at 5 turns with 2 model calls per turn is 400 calls plus judges before you have read a result, which is why the docs put Simulation on the per-merge golden set and Adaptive on a periodic sweep.

Red teaming

Red teaming is Promptfoo's strongest area and the reason OpenAI bought it. The plugin catalogue lists 157 plugins across brand, compliance and legal, dataset, security and access control, trust and safety, and custom categories, with presets mapped to the OWASP Top 10 for LLMs, NIST AI RMF, MITRE ATLAS and ISO 42001. The strategy list has 14 static single-turn strategies such as Base64 and Leetspeak, 12 dynamic single-turn strategies such as Best-of-N and tree-based attacks, and 5 multi-turn strategies, Crescendo, GOAT, Goblin, Hydra and Mischievous User. One caveat the docs state plainly is that some plugins and strategies, including most harmful content plugins and the Meta Agent and Best-of-N approaches, require remote inference in the Community edition, so a fully local red teaming run covers a subset of the catalogue. The Community plan also caps red teaming at 10,000 probes a month.

EvaliQA's red teaming datasets are built from the same two ingredients, a vulnerability and an attack technique, generated at a chosen number of attacks per vulnerability with a default of 3. The vulnerability catalogue is grouped into content safety, data and privacy, system behaviour, and access control for tool-using agents, where BFLA, BOLA, SQL injection and SSRF live. Techniques range from Direct and Authority through Roleplay and Hypothetical to Base64, ROT13 and embedded JSON, and multi-turn plans add Crescendo, Linear jailbreak, Persistent context, Context poisoning and Context flooding, with an escalation style of Gradual, Aggressive or Stealth. A severity filter lets you run only critical rows on every pull request and the full range weekly. Red teaming is a Team plan feature in EvaliQA, so it is not free, which is a fair point in Promptfoo's favour for a solo developer.

Where EvaliQA pulls ahead is what happens after the run. A red teaming run skips the metrics step entirely because the verdict is scored per vulnerability and technique, the results page breaks breaches down by category, and the generated report carries a per-vulnerability breach summary that a product manager can read. The docs also give targets, a breach rate of 0 percent on content safety and access control, under 5 percent on prompt injection and jailbreak for a well-prompted system, under 1 percent on data exfiltration, and the instruction that a miss on those is an incident rather than a prompt tweak. Promptfoo's plugin count is larger, and if your need is framework compliance coverage, that count is a real advantage. If your need is a red teaming pass that runs on a schedule, compares against last month and lands in front of the people who decide to ship, the platform matters more than the catalogue.

Production monitoring and the loop back

This is the section where the two tools stop overlapping. Promptfoo's tracing is built on OpenTelemetry with a built-in OTLP receiver, and its stated purpose is to show what the application did behind each response during an evaluation run, with visibility into tool calls and RAG steps and the ability to pull spans from Grafana Tempo, Braintrust or Langfuse. It is evaluation-time tracing. The docs do not describe scoring live user traffic, and the Enterprise tier lists continuous monitoring as a feature without detail, so we will not characterise it beyond that.

EvaliQA runs two loops on one metric library. Your agent reports traces through the Python SDK or any OTLP/HTTP exporter, they appear within seconds grouped into sessions and users, and the project's Observability tab aggregates volume, error rate, cost and p95 latency. With online evaluation switched on, every bound metric scores each trace as it arrives, and 60 seconds after a session's last trace the conversation metrics, the deterministic detectors and, if you want it, the AI Analyst run over the whole dialogue. The detectors cost nothing and catch explicit errors, retry loops, the same tool called twice with the same arguments, a session that reaches 0.50 US dollars, and a trace three times slower than the session's median. Alert rules watch pass rates, mean scores and finding severities and post to Slack, Discord, Teams, Telegram, email or a webhook, throttled so a bad deploy is one message.

The reason this belongs in an evaluation comparison and not an observability one is the last step. The trace behind an alert becomes a dataset row with Add trace to dataset, and the next run covers it. That is the loop a config-file CLI cannot close on its own, because the test cases and the production traffic live in different systems with a human copying between them. Without it the golden set stagnates and you keep scoring the same 60 rows while production finds failure modes you never test. The limit is equally plain. A trace without a session id is stored but never analysed, history is not re-scored when you change a binding, and every judge call on live traffic is a model call on your key, so the docs tell you to bind the metrics you would act on rather than the whole catalogue.

Comparing runs and reporting

Promptfoo's answer to "did the change help" is the matrix, and for a prompt change it is a good answer because both versions are in the same grid. The web UI shows outputs across prompts and providers with assertion results, and for red teaming it produces a vulnerability and risk report. What the Community edition does not give you is a shared history that outlives a laptop, since team sharing and a central dashboard are Enterprise features.

In EvaliQA the Compare view lives on the test plan page. Open the plan, switch to the Eval runs tab, tick 2 or more runs and press Compare, and you get per-metric pass rates, cost and latency side by side plus the list of rows that flipped verdict. The docs lay out 4 kinds of comparison, a prompt or model tweak, a dataset refresh, a model swap and a metric change, and what to hold constant for each, which is the part people get wrong. A dataset refresh that adds 50 hard rows will drop the pass rate, and that is coverage gained rather than a regression. Compare only works within one plan, and if you want to compare across plans the docs send you to the CSV export, which is a limit worth knowing before you design your plans.

Reporting is where the audience changes. From any completed run, Generate report drafts an executive summary, headline numbers, a per-metric breakdown, notable failure patterns by category and persona, a red teaming breach summary where relevant, and suggested next steps, in 15 to 60 seconds, editable inline and exportable to PDF, DOCX, HTML or Markdown. The docs are careful to call it a first draft. The sentence a product manager needs, whether this ships or not, is yours to write, and the draft can mis-state a number, so you read it against the run before sharing. Promptfoo has no equivalent in the open-source edition, and for a team whose evaluation results need to reach someone who will never open a terminal, that is the single most practical difference in this post.

Cost, licence and ownership

Promptfoo's Community edition is free with all evaluation features, all providers, local or self-hosted deployment and red teaming capped at 10,000 probes a month, with Enterprise and On-Premise tiers on quote adding team sharing, continuous monitoring, SSO, API access, managed cloud and a dedicated runner. EvaliQA's Free plan is permanent rather than a trial, with 2 projects, 2 seats, 1,000 evaluated cases and 5,000 runtime traces a month. Team is 99 US dollars a month or 948 a year for 10 projects, 5 seats, 10,000 cases, 25,000 traces, voice evaluation and red teaming. Custom is unmetered and adds self-hosted deployment, SSO and SAML, extended roles, audit log export and a signed data processing agreement. On both tools you pay your model provider directly for tokens. EvaliQA also gives each new account 5,000 welcome credits so a first run needs no key at all.

Dimension

Promptfoo

EvaliQA

Form

Open-source CLI and library, MIT licence, local web UI

SaaS or self-hosted on Custom plan

Test cases

Declared in a config JSON file or loaded from files

Hand-written, generated, imported or promoted from traces, versioned

Judge

Provider per assertion

One judge per test plan, pinned

Multi-turn

Simulated user provider, 10 turns by default

Simulation, Scripted and Adaptive strategies, 15 personas

Red teaming

31 strategies including 5 multi-turn, some need remote inference

Vulnerability and technique catalogue, 3 escalation styles

Production

Tracing scoped to evaluation runs

Live traces, online evaluation, detectors, alerts

Reporting

Matrix view and red team risk report

Compare view, generated executive and detailed report exportable to PDF, DOCX, HTML, Markdown

Pricing

Free Community edition, Enterprise on quote

Free, Team at 99 US dollars a month, Custom on quote

Ownership is the quieter question. Promptfoo is now part of OpenAI, and the announcement commits to keeping it open source and supporting a diverse range of providers, with the core technology to be integrated into OpenAI's model and infrastructure layers. We take that commitment at face value. It is still reasonable for a team that evaluates Anthropic, Google or open models to watch where the roadmap goes, and the MIT licence means the code you have today stays yours either way. EvaliQA is an independent vendor whose business depends on scoring whichever model you point it at, which is a different kind of guarantee with its own trade-off, namely that you are relying on a hosted service rather than a binary on disk. Data retention follows the plan, 14 days of traces and 30 days of row-level results on Free, 180 and 365 on Team, with summary scores, datasets and reports kept on every plan.

If you are deciding, run the same 30 cases through both. Put them in a Promptfoo config with an llm-rubric assertion and in an EvaliQA single-turn plan with Answer Relevancy and one custom metric, using the same judge model in both, and read the per-row reasoning on the 5 rows the two tools disagree about. That hour tells you more than any comparison post, including this one. If Promptfoo's matrix is the view you keep opening, your work is prompt iteration and Promptfoo is the right tool for it.

If instead you find yourself exporting results to share them, copying production failures into the config by hand, or being asked by a product manager what the numbers mean, you have hit the edge this post is about. Start with Quick eval against your endpoint, which creates the project, connector, plan, dataset and run from one page. Then follow the evaluation pipeline one rung at a time, a 20 to 50 row golden set gating pull requests through the CI endpoint, a nightly run, then the SDK in production with online evaluation on 3 cheap metrics, and the weekly habit of promoting flagged traces into the dataset. Wire one rung fully before adding the next. A half-wired gate that gets ignored is worse than no gate, whichever tool is behind it.

Frequently asked questions

Can I use EvaliQA without writing code?

Yes for everything up to and including the first evaluation run. Quick eval takes a pasted cURL command and one or two sentences about what to test and creates the project, connector, test plan, dataset and run from one page. The Python SDK or an OpenTelemetry exporter is only needed to stream production traces.

Which tool is better for red teaming?

Promptfoo has catalogue with presets for OWASP, NIST, MITRE ATLAS and ISO 42001, and 5 multi-turn strategies, though some plugins and strategies need remote inference in the Community edition. EvaliQA generates attacks from a vulnerability and technique catalogue with Gradual, Aggressive or Stealth escalation, scores breaches per vulnerability, and carries the result into a scheduled cadence, a run comparison and a generated report. Red teaming in EvaliQA is a Team plan feature.

Can I move my Promptfoo test cases into EvaliQA?

The inputs and expected outputs, yes. Export them to CSV or JSONL and upload the file from the New Dataset sheet on the Datasets page. A CSV needs a header row, and the header names become the dataset columns; a JSONL file maps object keys to columns the same way. The limit is 1,000 rows per file, and every imported column arrives as text, so retype numbers and enums from the column menu before you run.

The assertions do not carry over, because in EvaliQA a check belongs to the test plan rather than to the row. Rebuild them as metrics on the plan, deterministic ones such as Regex Match, JSON Schema, Length Check and Contains for the mechanical checks, and a custom metric written in plain language for anything that was an llm-rubric. A javascript or python assertion has no direct equivalent and has to be expressed as a rubric.

How do I run EvaliQA in CI the way I run Promptfoo?

Create an API token on the Integrations page with the workspace scope that includes runs create, store it in your CI secrets, and call the trigger endpoint when a pull request opens. The call returns a run id; the job either polls until the run is completed or waits for a webhook callback carrying the pass rate, and merge is gated on a pass rate threshold. The same endpoint on a cron gives you the nightly run, and if scheduled runs are enabled in your workspace the schedule can be set on the plan page instead.

Trigger the plan that finishes in under a minute, a single-turn golden set of 20 to 50 rows. Multi-turn and red teaming plans belong on the nightly or weekly schedule, not on every pull request. Leave room for judge noise in the threshold, because a 3 percentage point drop can be the judge rather than your change.

Do I need my own model API key to try EvaliQA?

No. Every workspace starts with 5,000 bonus credits, and if you pick them at sign-up the judge, dataset generation and every other AI operation run on the platform model with nothing to paste. A typical judge call costs tens of credits and a 50-row run with 4 metrics a few thousand, so the welcome balance covers a first run and a little more. When the credits run out a platform-model call is refused with a No credits left message, and you either buy a pack or switch the Platform AI agent to your own key under Settings, Platform AI agent.

On your own key EvaliQA never bills for tokens. Model calls land on your provider's invoice, and any of the 100 plus supported providers can be stored as a credential, including a local Ollama server that needs only a URL.

Get new posts by email

One email when something new is published. No spam, unsubscribe any time.