the short answer
As checked September 22, 2026, TypeSafe documents Jev’s typed primitives, current model aliases, direct price, limits, English-first guidance and no-training use of requests, and announced general availability without a waitlist. Its latency and benchmark statements remain first-party measurements. LangChain provides one independent evaluator experiment, while public field reports include both working prototypes and a negative reranking result. Universal superiority remains unverified.
- Last verified
- September 22, 2026
- Official facts
- Interface, model listing, pricing, limits and stated data use
- First-party measurements
- Launch latency and benchmark claims
- Independent evidence
- LangChain evaluator report and attributed field reports
- Not established
- Universal superiority or calibration across tasks
Current Claim and Evidence Ledger
| Claim | Evidence class | Status |
|---|---|---|
| Jev returns Choice, Score and Noul decisions | Official docs | Documented |
| Jev is generally available with no waitlist | TypeSafe first-party X announcement | Announced; recheck the current access flow |
| jev-1.13.0 is current stable; aliases resolve to it | Official model page checked 2026-09-22 | Dated fact |
| $0.042 per million direct input tokens; output free | Official model page checked 2026-09-22 | Dated direct price |
| 70–500 ms response time | TypeSafe launch measurements | First-party, environment-dependent |
| Requests and responses are not used for training | Official model page | Documented; not a retention claim |
| Jev is generally better than LLM judges | No broad independent benchmark | Not established |
Different Sources Prove Different Things
Citations should preserve this level in the sentence. “TypeSafe reports 70–500 ms” is supportable; “Jev responds in 70 ms” removes both the range and attribution. “A project reports X” is not the same as “Jev achieves X.” The social evidence roundup applies this rule to launch commentary and community experiments.
| Evidence | Can establish | Cannot establish alone |
|---|---|---|
| Official documentation | Current contract, listed limits, price and declared data handling | Independent performance or fitness for your task |
| Vendor benchmark | Result for the vendor’s disclosed setup | Universal superiority or your production latency |
| Independent experiment | Result for a named external dataset and configuration | Transfer to a different criterion or population |
| Repository or demo | An implementation artifact exists | Security, maturity, production adoption or general quality |
| Social post | An attributable claim or discovery lead | Reproducibility when methods and outputs are absent |
Latency Claims Need Workload and Network Context
TypeSafe reports a 70–500 ms end-to-end range and says its published evaluations were generally run from West Coast laptops near the service. Preserve that attribution and context. State length, question count, provider route, region, load, retries and client work can change observed latency.
Measure p50, p95 and p99 end to end in the deployed architecture. Do not shorten a first-party range into “70 ms latency” or compare it with model-only timing from another system.
Price Facts Require a Route and Date
TypeSafe’s direct model page listed $42 per billion input tokens—$0.042 per million—and free output tokens when checked. Provider pricing may differ, and retries, gateways, storage and human review add cost. The pricing guide gives worked examples.
Rate-limit figures are explicitly dynamic. Treat 250k tokens per second and 1,200 requests per minute as dated listed limits, not contractual capacity.
Calibration Is a Training Objective, Not a Deployment Guarantee
TypeSafe describes RLCD as reinforcement learning for calibrated decisions. That supports the design objective, not the claim that every probability is calibrated for every criterion or population. Validate reliability bins, Brier or log loss and action coverage on target labels.
LangChain’s evaluator experiment is independent evidence for one configuration. The reranking post is a useful negative result for another. Together they argue for task-specific testing, not an overall winner.
Availability Claims Need Both Announcement and Interface Evidence
TypeSafe’s cited September 2026 post says Jev is generally available with no waitlist. That answers the broad access question at that date. A developer should still verify the current sign-up flow, supported region, credentials and account limits before planning a deployment.
Provider announcements need an additional check. Cloudflare’s announcement is backed by a published model page and request contract. OpenRouter announced Jev beta, but its public models catalog did not expose a TypeSafe/Jev entry when checked September 22, 2026. The accurate statement includes both observations rather than selecting the more convenient one. See the platform comparison and launch tracker.
No-Training and No-Retention Are Different Claims
The official model page states that requests and responses are not used for training. Retention duration, regional processing, abuse logs, backups and third-party routes require separate evidence. See the privacy and retention review.
Maintain the Page as a Dated Correction Record
- Link the primary source beside each claim.
- Record verification date, provider and resolved model.
- Label vendor, independent, community and anecdotal evidence.
- Preserve prior wording when a factual correction materially changes a conclusion.
- Downgrade stale claims to unknown rather than carrying them forward.
- Invite reproducible counterevidence and link the artifact.
FAQ
Is Jev proven to be better than LLM judges?
No universal conclusion is supported. Compare on the target task with blind labels and pinned configurations.
Is Jev always 70 ms?
No. TypeSafe reported a broader 70–500 ms range under described conditions. Measure your complete route.
Is the listed price permanent?
No. It is a dated direct-provider price and should be rechecked before budgeting.
Does RLCD guarantee calibrated scores?
No. It is TypeSafe’s stated optimization objective. Calibration must be measured on the deployed criterion and population.
Is Jev still on a waitlist?
TypeSafe announced general availability without a waitlist in September 2026. Verify the current official access flow before implementation.
Sources
Checked against the sources below on September 22, 2026. Model versions, prices and limits change.
- TypeSafe AI: Introducing System One Models and Jev
- TypeSafe AI docs: Models
- TypeSafe AI docs: Primitives
- TypeSafe AI docs: Jev 1.13 jaggedness
- LangChain: Can Jev be a better agent evaluator?
- GoSailGlobal on X: Jev reranking field report
- TypeSafe AI on X: Jev general availability
- Cloudflare Developers on X: Jev gateway launch
- OpenRouter on X: Jev beta announcement