the short answer
InternLM publishes Intern-Decision as Apache 2.0 checkpoints at 0.8B, 2B and 4B, with local inference code and optional image input. Jev is TypeSafe’s separate hosted model. InternLM’s benchmark table is a publisher-run result on a selected panel. If you need image-assisted triage or local serving, run both paths on matched labeled cases, then calibrate and set thresholds for each checkpoint separately.
- Publisher
- InternLM; independent of TypeSafe AI
- Models
- Intern-Decision-0.8B, 2B and 4B
- Artifact
- Apache 2.0 model cards, weights and local inference source
- Input
- Shared state and typed questions; model card documents optional images
- Evidence
- InternLM-run benchmark and calibration tables; target-task replication needed
Inside InternLM's Answer Scoring
The 4B card says Intern-Decision fine-tunes Qwen3.5-4B and scores the allowed answers for each named question. It also lists 0.8B and 2B checkpoints. Its inference compiler builds a complete answer skeleton and reads token scores where decisions appear. Those details describe InternLM’s model and local code.
Jev has its own Choice, Score and Noul contract. You may be able to adapt the same question definitions, but compare response fields and probability behavior before sharing application logic. The InternLM card documents a Python DecisionEngine, not a hosted endpoint.
Compare an Image-Assisted Triage Task
Give Intern-Decision the screenshot and you have tested a visual system. Jev’s documented path is text, so a paired text test needs a reviewed transcription of the error banner. Keep the transcription errors in that result. Otherwise the score would hide which system saw the crucial words.
| Field | Matched experiment |
|---|---|
| State | A support message and screenshot of an error banner |
| Choice | Which team resolves the first blocker: access, billing, technical or review? |
| Noul | Does the screenshot itself show a payment failure? |
| Text baseline | Repeat with a reviewed text transcription so Jev sees comparable evidence |
| Record | Native output, checkpoint, image or transcription, reviewer label and error |
Read Benchmark Results at Their Stated Scope
InternLM tested Jev and three Intern-Decision checkpoints on its chosen typed-decision, safety and classification tasks. It also reports latency. Those numbers depend on its harness, hardware and calibration, so inspect the task rows before carrying any claim into a ticket or policy workflow.
One calibration pilot scores distance from exact reference distributions. If your production label is a yes/no outcome, that is a different target. Pin the revision, temperature and source code before reproducing the pilot or fitting probabilities to your own outcomes.
How to Test Substitution
Save the exact checkpoint and every native result in the benchmark record. The open-model guide covers the separate weight, license and wire-contract checks.
- Freeze one task, answer definitions, evidence projection and blind labels.
- Pin Jev model and provider plus the exact Intern-Decision checkpoint and local code.
- Check request and response behavior for every primitive, invalid input and timeout.
- Compare paired errors, per-slice calibration and review coverage; fit thresholds independently.
- Measure the complete hosted or local path, including GPU capacity, failures and review work.
FAQ
Is Intern-Decision an official Jev model?
No. InternLM publishes it independently of TypeSafe AI.
Can Intern-Decision run locally?
The model cards publish weights and Python inference code. Verify the exact checkpoint, license, hardware and dependencies before deployment.
Does Intern-Decision beat Jev?
Its publisher reports results on a selected panel. Only a paired evaluation on your labels and operating point can answer that for your workflow.
Sources
Checked against the sources below on October 2, 2026. Model versions, prices and limits change.
- InternLM: Intern-Decision-4B model card
- InternLM: Intern-Decision source
- TypeSafe AI: Jev primitives
- TypeSafe AI: Jev models
Change note: Added a source-checked comparison of InternLM’s open decision-model family.