Jev knowledge base·verified October 2, 2026

jev evals in deepeval

Use DeepEval JevEval and ConversationalJevEval for bounded questions, and understand score calculation, strict mode and probability limits.

the short answer

DeepEval provides JevEval for single-turn tests and ConversationalJevEval for conversations. Pick the test-case fields Jev can see, write Noul, Choice or Score questions, and DeepEval turns the returned distributions into a weighted score. Its text reason is assembled in code. That score reflects your questions and weights; it is not a measured chance of agent success. Check it against blind human labels before setting a release threshold.

Single turn
JevEval over selected SingleTurnParams fields
Multi-turn
ConversationalJevEval over selected MultiTurnParams fields
Questions
Noul, Choice and Score with caller-defined weights
Output
Weighted metric score, breakdown, threshold result and deterministic text reason
Verification
DeepEval docs checked October 2, 2026; live SDK call not run

Choose the Fields Before the Questions

DeepEval’s JevEval docs turn selected test-case fields into Jev state. You supply typed questions, option credits and weights. The wrapper calculates the metric score. ConversationalJevEval does this for a conversation. DeepEval also describes using Jev inside other metrics in an earlier integration post, which is a separate path.

A question about tool-result faithfulness needs the actual tool result in state. If you only pass the final answer, Jev cannot check its source. The trace-state guide shows how to select that evidence without dumping an entire unbounded run into one request.

A Bounded Tool-Result Faithfulness Metric

The imports and field names follow DeepEval’s documentation. Check them against your installed package. With threshold=None, you can collect scores and labels before choosing a gate. This example leaves claim splitting and tool-result selection to the caller. For a long answer, use the groundedness guide to test individual claim-evidence pairs.

from deepeval.metrics import JevEval
from deepeval.metrics.jev_eval import Noul, Score
from deepeval.test_case import SingleTurnParams

tool_faithfulness = JevEval(
    name="Tool faithfulness",
    evaluation_params=[
        SingleTurnParams.ACTUAL_OUTPUT,
        SingleTurnParams.TOOLS_CALLED,
    ],
    questions=[
        Noul("Every factual claim in actual_output is supported by tools_called."),
        Score("How much of actual_output is grounded in tools_called?", levels=[
            "Fabricated", "Mostly fabricated", "Mostly grounded", "Fully grounded"
        ]),
    ],
    threshold=None,
)

Read the Metric Score Without Changing Its Meaning

For a Choice answer marked with None credit, DeepEval drops the question when those answers take at least half the probability. Look at score_breakdown when the aggregate rises. The metric may have skipped a difficult question rather than seen a better agent run.

DeepEval builds its text reason from the decisions in code. It is a readable record of the metric’s own outputs, with no new citation or Jev-generated explanation. Check the judgments against labels using evaluator validation.

ValueWhat it representsWhat it does not prove
Jev Noul probabilityModel support for one defined yes/no criterionEmpirical correctness for your traffic without calibration
Choice or Score mappingCredit assigned by the metric author to an answer or levelA natural probability of agent success
JevEval weighted scoreWeighted mean of applicable question valuesA calibrated chance that the whole run succeeded
DeepEval pass resultScore compared with the configured thresholdPermission to automate a consequential action

Choose Score-Only, Thresholded or Strict Use Deliberately

The default 0.5 threshold ships with the software. It has not been fitted to your false-pass costs. Choose a threshold on a calibration split, then report false passes, false failures and review load on untouched cases. The threshold guide covers that step.

ModeDocumented behaviorBest starting use
Score-onlythreshold=None returns a score without a pass gateExplore distributions and collect labels
ThresholdedDefault threshold is 0.5; configured score determines passGate only after fitting a task-specific threshold
StrictRequires every applicable question to receive its best credit; threshold becomes 1Use only when every condition is genuinely required

Where JevEval Fits Among Eval Workflows

JevEval works when you can name the bounded questions and accept a weighted score as the result. A task that needs open-ended critique or evidence citations calls for a generative judge. The platform comparison covers what DeepEval, LangSmith, jevals and Failproof store around those judgments.

To replace a judge, replay the same blind cases through both and save the native outputs. Compare wrong calls at a chosen review rate, failed calls, latency and total cost. Check calibration where each result supplies a meaningful probability.

FAQ

Does JevEval require a generative LLM?

The documented custom JevEval metric uses Jev for its bounded questions and builds its reason in code. Other DeepEval metrics can still use an LLM for their own steps.

Is a JevEval score a probability of success?

It is a weighted metric score built from your questions, option credits and applicability rules. Compare score bands with observed outcomes before reading it as a success probability.

Can JevEval evaluate conversations?

DeepEval documents ConversationalJevEval for multi-turn test cases. Choose fields and questions that refer to the whole conversation or a named role.

Is strict mode safer?

Strict mode requires maximum credit on every applicable question. It can still pass bad runs or reject good ones, so count both errors on a held-out set before using it as a gate.

Sources

Checked against the sources below on October 2, 2026. Model versions, prices and limits change.

  1. DeepEval: JevEval documentation
  2. DeepEval: Introducing JevEval
  3. DeepEval: Introducing Jev for evals
  4. TypeSafe AI: Jev primitives
  5. TypeSafe AI: Confidence

Change note: Added a source-checked guide to DeepEval’s published JevEval integration.