The Playground and Writing Good Questions
The playground lets you put a question to Jev about sample text before any feature depends on it. It is where you find out whether a question is worded well enough to trust, and what confidence bar it needs.
Every playground call is real, billed and logged under the feature jev.playground.
Where to find it
Architect Panel → Activity:
- TypeSafe Jev — the Playground card, below Check the connection
The card is missing when Playground on the Jev console is switched off in AI Settings.
Asking a question
- Write the state. This is the text or data the questions are about, such as a customer message. Choose Text, or JSON (an object or an array) to send several labelled values. The counter shows how much of the 20,000-character limit you have used.
- Choose the model. Default uses the installation's model. You can also pick
jev-latest,jev-previewor A pinned version and type one such asjev-1.13.0. - Set up each question. Give it an Id (letters, digits,
_ . : -, up to 64 characters), a Type and its Instructions, which say what Jev is being asked.- Choice (pick one): under Options, one per line as
key: description. A colon followed by a space separates the two, so keys such as09:00orGB:VAT20keep their own colons. The description is optional; up to 100 options. - Score (ordered levels): under Levels, one per line, lowest first, 2 to 10 of them.
- Yes / no: the instructions are the statement. What "yes" means and What "no" means are optional.
- Choice (pick one): under Options, one per line as
- Add a question for more, up to 20 in one call, or Remove one.
- Click Ask Jev.
If anything is wrong, such as a duplicate option or a score with one level, the result says Not sent and lists the problems. Nothing reached TypeSafe, so it cost nothing and was not logged.
Reading the result
- Each answer shows a bar and a percentage for every option or level, with the winner highlighted. A choice gives its confidence and its margin over the runner-up; a score gives the score and the most likely level; a yes/no gives the probability that the statement is true.
- The outcome badge (confident, review, yes or no) uses the starting bars: 0.6 confidence for a choice or score, and 0.75 and 0.25 for yes and no. A sentence underneath explains it.
- The call details: the model that answered, TypeSafe's request id (quote it to TypeSafe if you need their help), tokens, cost, attempts, duration and the run, linked to AI Usage.
Writing good questions
- Send only what the question needs. Accuracy falls as unrelated detail grows. A subject and a description beat a whole record.
- One narrow judgement per question. Ask "which team" and "is it urgent" separately, then combine the answers in a workflow.
- Say exactly what you mean. Jev reads literally. Put the boundary cases in the option descriptions ("Payments, refunds and invoices; not pricing questions"), and add a none of these option whenever nothing may fit.
- Keep arithmetic out. Counting, sums and date comparisons belong in a Condition step or a rule. Ask Jev for the judgement only.
- Assume the text was written by a stranger. A message can claim to be urgent or argue for its own category. Test examples that try.
Calibrating the bar
Run a dozen real examples through the same question, including the awkward ones, and note the confidence each time. Where right answers cluster above a figure and wrong ones fall below it, that figure is your bar. Write the model version down beside it: a bar tuned against one version may not hold on the next, which is why the Model setting can be pinned.
Taking it into a feature
- Workflow AI decision: the instructions, options and levels go into the step's Instructions, Choose from and Levels, lowest first. The bar goes into Sure enough at.
- Workflow AI check: the statement goes into Statement to check, with Yes means and No means.
- Is about (AI) message rule: test the statement as a yes/no question with a real message as a text state. Rules allow up to 500 characters.
Worked example
An operations lead wants a workflow to route customer enquiries by intent. In the playground they paste a real message, add a choice question with exchange, refund, track and other, and ask Jev. It picks exchange at 91%. They try twelve more messages; the two it got wrong both scored under 70%, so they set the workflow step's Sure enough at to 0.75 and pin the model the console reported.
Recommendations
- Use real but harmless examples. The playground sends what you type to TypeSafe, so leave out names and account numbers.
- Test the edge cases, not just the easy ones.
- Keep the wording you settle on and copy it exactly into the feature.
- Re-test after a model change before trusting an old bar.