the short answer
DeepEval provides JevEval for single-turn tests and ConversationalJevEval for conversations. Pick the test-case fields Jev can see, write Noul, Choice or Score questions, and DeepEval turns the returned distributions into a weighted score. Its text reason is assembled in code. That score reflects your questions and weights; it is not a measured chance of agent success. Check it against blind human labels before setting a release threshold.
- Single turn
JevEvalover selectedSingleTurnParamsfields- Multi-turn
ConversationalJevEvalover selectedMultiTurnParamsfields- Questions
- Noul, Choice and Score with caller-defined weights
- Output
- Weighted metric score, breakdown, threshold result and deterministic text reason
- Verification
- DeepEval docs checked October 2, 2026; live SDK call not run
Choose the Fields Before the Questions
DeepEval’s JevEval docs turn selected test-case fields into Jev state. You supply typed questions, option credits and weights. The wrapper calculates the metric score. ConversationalJevEval does this for a conversation. DeepEval also describes using Jev inside other metrics in an earlier integration post, which is a separate path.
A question about tool-result faithfulness needs the actual tool result in state. If you only pass the final answer, Jev cannot check its source. The trace-state guide shows how to select that evidence without dumping an entire unbounded run into one request.
A Bounded Tool-Result Faithfulness Metric
The imports and field names follow DeepEval’s documentation. Check them against your installed package. With threshold=None, you can collect scores and labels before choosing a gate. This example leaves claim splitting and tool-result selection to the caller. For a long answer, use the groundedness guide to test individual claim-evidence pairs.
from deepeval.metrics import JevEval
from deepeval.metrics.jev_eval import Noul, Score
from deepeval.test_case import SingleTurnParams
tool_faithfulness = JevEval(
name="Tool faithfulness",
evaluation_params=[
SingleTurnParams.ACTUAL_OUTPUT,
SingleTurnParams.TOOLS_CALLED,
],
questions=[
Noul("Every factual claim in actual_output is supported by tools_called."),
Score("How much of actual_output is grounded in tools_called?", levels=[
"Fabricated", "Mostly fabricated", "Mostly grounded", "Fully grounded"
]),
],
threshold=None,
)Read the Metric Score Without Changing Its Meaning
For a Choice answer marked with None credit, DeepEval drops the question when those answers take at least half the probability. Look at score_breakdown when the aggregate rises. The metric may have skipped a difficult question rather than seen a better agent run.
DeepEval builds its text reason from the decisions in code. It is a readable record of the metric’s own outputs, with no new citation or Jev-generated explanation. Check the judgments against labels using evaluator validation.
| Value | What it represents | What it does not prove |
|---|---|---|
| Jev Noul probability | Model support for one defined yes/no criterion | Empirical correctness for your traffic without calibration |
| Choice or Score mapping | Credit assigned by the metric author to an answer or level | A natural probability of agent success |
| JevEval weighted score | Weighted mean of applicable question values | A calibrated chance that the whole run succeeded |
| DeepEval pass result | Score compared with the configured threshold | Permission to automate a consequential action |
Choose Score-Only, Thresholded or Strict Use Deliberately
The default 0.5 threshold ships with the software. It has not been fitted to your false-pass costs. Choose a threshold on a calibration split, then report false passes, false failures and review load on untouched cases. The threshold guide covers that step.
| Mode | Documented behavior | Best starting use |
|---|---|---|
| Score-only | threshold=None returns a score without a pass gate | Explore distributions and collect labels |
| Thresholded | Default threshold is 0.5; configured score determines pass | Gate only after fitting a task-specific threshold |
| Strict | Requires every applicable question to receive its best credit; threshold becomes 1 | Use only when every condition is genuinely required |
Where JevEval Fits Among Eval Workflows
JevEval works when you can name the bounded questions and accept a weighted score as the result. A task that needs open-ended critique or evidence citations calls for a generative judge. The platform comparison covers what DeepEval, LangSmith, jevals and Failproof store around those judgments.
To replace a judge, replay the same blind cases through both and save the native outputs. Compare wrong calls at a chosen review rate, failed calls, latency and total cost. Check calibration where each result supplies a meaningful probability.
FAQ
Does JevEval require a generative LLM?
The documented custom JevEval metric uses Jev for its bounded questions and builds its reason in code. Other DeepEval metrics can still use an LLM for their own steps.
Is a JevEval score a probability of success?
It is a weighted metric score built from your questions, option credits and applicability rules. Compare score bands with observed outcomes before reading it as a success probability.
Can JevEval evaluate conversations?
DeepEval documents ConversationalJevEval for multi-turn test cases. Choose fields and questions that refer to the whole conversation or a named role.
Is strict mode safer?
Strict mode requires maximum credit on every applicable question. It can still pass bad runs or reject good ones, so count both errors on a held-out set before using it as a gate.
Sources
Checked against the sources below on October 2, 2026. Model versions, prices and limits change.
- DeepEval: JevEval documentation
- DeepEval: Introducing JevEval
- DeepEval: Introducing Jev for evals
- TypeSafe AI: Jev primitives
- TypeSafe AI: Confidence
Change note: Added a source-checked guide to DeepEval’s published JevEval integration.