Choosing Jev or Laya? See 5 tests before swapping your LLM for a decision model
Want to evaluate quality, confidence, and cost before putting Jev or Laya in the path of an automation? Let's go! Classifying tickets and routing requests can be quite simple!

Want to evaluate quality, confidence, and cost before putting Jev or Laya in the path of an automation? Let's go! Classifying tickets and routing requests can be quite simple!
Imagine you received a message like this in support: “The payment went through fine, but my access is still blocked. And I need this as soon as possible for a meeting coming up soon.”
For a person, the problem is completely understandable. But for an automation, it requires decisions: route to billing or technical support? Mark as urgent? Ask for more information?
It's in this space that Jev and Laya spark interest. Instead of producing a long response and then extracting a classification, the proposal is to deliver structured decisions that software can consume.
iMasters has already introduced Jev and use case examples and covered Laya as an open alternative. For those planning to adopt one of them, the next step is to build an evaluation that represents real work.
The roadmap below proposes five tests. It doesn't assume that one tool is superior to the other: the goal is to find out where each option can take on a task with measurable benefit.
1. Does the tool understand your problem, including in Portuguese?
Start with a specific decision. “Automating support” is too broad. “Routing requests between billing, technical support, and sales” allows you to define what a correct response would look like.
Build a sample of real messages, with personal data removed when not necessary. Include short texts, typos, negations, abbreviations, and requests that mix subjects.
These examples illustrate situations that deserve to be part of the test:
| Message | What it requires from the classification |
|---|---|
| “I don't want to cancel, just change the card.” | Distinguishing a mention of cancellation from an intent to cancel |
| “Pix approved, access blocked.” | Applying a clear rule between billing and support |
| “You can update it, but only after 10 PM.” | Preserving the condition set by the user |
| “It messed up again.” | Recognizing that context is missing |
Before running the models, ask the team to define the expected outcomes. If two people disagree about the correct queue, perhaps the first adjustment needed is in the support rule itself.
Compare Jev and Laya with the current system, whether it's an LLM, a classifier, or simple rules. Use the same cases and criteria. Set aside a portion of the examples for the final evaluation, without using it to adjust prompts, training, or confidence thresholds.
2. Are you measuring probability, or just reading a confidence number?
A numerical output is useful, but it requires interpretation.
In Jev's documentation, Choice and Score return a probability distribution and a confidence field, calculated from that distribution. Noul, on the other hand, returns a value between zero and one, without that additional field. Therefore, you shouldn't automatically treat every confidence: 0.9 as "a 90% chance of being right." TypeSafe's confidence documentation.
For class probabilities, a practical question is: among predictions close to 90%, does the accuracy rate also stay close to that? This correspondence is the central idea behind calibration. Research shows that neural networks can present poorly calibrated probabilities, even when they classify well. Study on neural network calibration.
In the project, log the returned number, the decision, and the correct outcome. Analyze by category and input type. An overall average can hide poor behavior precisely in the most important messages.
The threshold used to automate should come from this evaluation. Copying 0.9 from a documentation example doesn't prove that value is right for your domain.

3. How many cases can be automated without increasing relevant errors?
Getting a lot right among a few selected answers is different from solving most of the operation.
Consider a hypothetical scenario: out of a thousand requests, the system automatically routes 400 and gets 392 right. The accuracy of automatic routing is 98%, but coverage is 40%. The other 600 requests still need handling.
Now imagine a configuration that automates 800 cases and gets 760 right. Accuracy drops to 95%, while coverage rises to 80%.
Which configuration is better? It depends on the cost of errors and the capacity for review.
Routing a ticket to the wrong team can create rework. Failing to identify an urgent incident can have a much bigger consequence. That's why you should track separately:
- The share of automated cases.
- Errors among the automatic decisions.
- Critical cases that went unnoticed.
- The volume and time of human review.
In Laya's technical announcement, the author himself presents results that change according to the share of cases in which the model agrees to answer. This reinforces the importance of evaluating quality and coverage together. The published numbers come from the project, not from an independent test conducted by iMasters. Laya's technical publication.
4. Does speed hold up when the request travels through the entire system?
The model's execution time is just one part of the user's wait.
A useful evaluation starts when the application receives the request and ends when the decision is available for the next step. This interval includes network, data preparation, queueing, inference, and any retries.
Measure both the median and the 95th percentile (the time below which 95% of requests finish). Also test concurrency, long messages, and startup after a period of inactivity.
In Laya's case, the repository describes relevant differences between keeping models loaded and having to reload them when switching between inputs of different languages. It also states that the comparative numbers for Jev came from third parties, with differences in samples and questions. The public table serves as a starting point, not as a performance guarantee for your application. Laya's official repository.
For your own comparison, record versions, hardware, application region, input size, and number of questions per call. Without this, a lower number may simply reflect different conditions.
5. Does the savings hold up once you include operations and review?
An API charges for the service. A self-hosted installation requires resources to run. In both cases, the bill keeps growing after inference.
As a proposed evaluation, add up:
Total cost = inference or infrastructure + operations + human review + rework caused by errors.
Divide this total by the number of cases completed correctly within the expected timeframe. This measure brings the comparison closer to the outcome that matters to the product.
Also test the behavior when the model fails: timeout, invalid response, unavailability, or out-of-scope input. The system should have a defined destination for these cases, such as a review queue.
And keep classification and authorization separate. Recognizing that a customer requested a refund doesn't authorize executing the payment. Identifying an intent to approve doesn't eliminate conditions, access limits, or confirmations required by the process.
The first experiment can work without changing a single decision

One way to start is to run the model in parallel with the current flow, just logging its suggestions. Operations continue as before, while the team compares decisions, latency, and cost.
Define in advance which results would justify releasing a portion of the automation. Then, start with reversible actions, also track a sample of high-confidence decisions, and preserve a path back to the previous flow.
Jev and Laya can be evaluated as components of a specific task: routing, prioritizing, selecting, or verifying. The decision to adopt them becomes clearer when the team can demonstrate what work was automated, how many errors remained, and how much each useful outcome cost.
Translated from the Brazilian Portuguese original · Read the original





