Jev Usecase 1 - Ticket Triage

Jev
There’s been a new model making the headlines lately, which differs from the usual crop of large language models we’ve become accustomed to.
Jev, by TypeSafe , is in my understanding a fancy classifier. It understands the world in greater depth, more akin to LLMs than previous classifiers. Primary difference here is that Jev does not respond with words. It responds with probabilities.
It can answer three types of questions:
Noul: Yes or No, through a probability rating from0.0to1.0. Closer to 0- no, closer to 1- yes. You define how you want to handle the thresholds in your app logic.Choice: pick one from available choices, like a multiple-choice test. Returns an overall confidence score, plus a probability rating for each option.Score: a scale, with a maximum range of 10. This one is slightly confusing as you describe each step on that ladder and what it corresponds to. You cant just describe that 1 is bad and 10 is great, because Jev wont know what 5 means. So you have to explain each step on that ladder. It’s like Choice, but with some averaging added in.
It is being offered cheaply, and works very fast.
This gives us what is effectively a decision engine. There are no shortage of use cases, as twitter will show you. Here I would like to explore one particular use case.
I’ve spent most of my professional life working IT operations, where tickets rule the day. Triaging incoming work quickly - and accurately - is a pain for many teams. Using LLMs for triage and troubleshooting is nothing new, but a fast decision engine enables us to do the same thing more cheaply and faster than before.
Demonstrator
I’ve thrown together a demonstrator that generates a dataset and runs it through Jev and an LLM, then compiles the results into a table for us to interpret.
The app is comprised of 4 screens:
- Configuration surface
- a space to configure your inference providers
- specify two separate decision engines, any combination of LLMs or Jevs.
- Dataset generator
- generates tickets using an LLM provider and model you specify
- capable of generating varied tickets, within the domain of technical support
- Triage criteria
- Define a list of questions you want each ticket evaluated against.
- Decision engine results
- user-supplied data with
- side by side results and verdicts
You can find the repo below.
Export Options
I’ve added two export formats: .csv and a self-contained .html.
This keeps the results portable and makes it easier to show them to someone without having to compile and run the app.
I’ve generated a what I’m hoping is a sufficiently varied dataset. I’ve subsequently went through that dataset with a fast google/gemma-4-e2b local model and another time with openai/gpt-6-luna through OpenRouter.
Best if you view exports on a large screen, it produces large tables.
Interpreting Results
The generated dataset centers around technical IT support, since that is the area I’m most familiar with, but this approach generalizes to other domains.
I’m going to pick out a few things that stood out to me in the dataset I went through.
typesafe/jev-1.13.0 vs google/gemma-4-e2b
First up, let’s compare Jev vs locally-hosted Gemma 4 E2B with thinking enabled. This is a very light and fast variant of Gemma, which should make it ideal for use cases like these.
Jev 1.13 vs Gemma 4 E2B Side-by-side evaluation table. Launch ReportJevwhilst more conservative in its confidence estimates, is usually more correct or closer to the truth. It seems to understand more nuance in the context provided.#2e88Jev assumed this was a hardware problem because user labelled it as such, whereas Gemma4 inferred it must be a piece of software because of terminology used.#3401a very ambiguous one. Gemma4’s recommendation follows a more traditional route of going through the service desk, while Jev routes straight to the IAM teams.
typesafe/jev-1.13.0 vs openai/gpt-6-luna
Two more rounds, this time against GPT-6 Luna on medium.
Jev 1.13 vs GPT-6 Luna (Round 1) First round comparison. Launch Report Jev 1.13 vs GPT-6 Luna (Round 2) Second round comparison. Launch Report#b565Luna correctly surmised that the criticality should be medium (single user affected, albeit severely) and correctly routed it to Service Desk as the first port of call. It did make a mistake in classifying this as a service request rather than an incident. Jev on the other hand correctly classed it as an incident, but overstated the criticality and routed it to a specialist team instead of going through SD first.#b813Luna made the more sensible determination- single frustrated user, not a network-wide outage. Jev assessed the problem to be larger in scope.#95e3the inverse; Jev more sensibly judged that this should go through SD first before going through to a specialist.
typesafe/jev-1.13.0 vs typesafe/jev-1.13.0
Then I ran Jev vs Jev. Answers were mostly consistent, with a handful of disagreements and deviation. For now it shows that Jev is mostly consistent in its responses, but will deviate from time to time.
Jev 1.13 vs Jev 1.13 Consistency comparison. Launch Report#b21bis notable in that Jev incorrectly assumed multiple impacted users, even though the ticket strongly suggests only one person was affected. This implies there is some variability or randomness to how it evaluates things.
Generalizations:
- Jev was more consistent in its answers. ie re-run Jev on the same dataset, and you will get mostly same answers. LLM will guess all over the place with every run.
- LLM were consistently confident in its answers, even when they are wrong. Confidently wrong, assigning a high confidence % for their answers. That’s not a surprise and to be expected from LLM.
- Jev probabilities are more varied and appear to be a more accurate reflection. This should probably be put to a test by an actual statistician.