Independent research notebook / vol. 01
Jev classification, tested in the open.
Decivo documents a narrow benchmark of Jev, TypeSafe AI's typed decision model, on a fixed customer-support sample. The study is in preparation; results are pending.
00 / premise
A narrow benchmark for typed decisions.
Jev is designed to return typed decisions from unstructured input. Decivo asks one bounded question: how does it classify 50 customer-support messages against one comparison model when the labels and edge cases are stated up front?
01 / method
Make each decision inspectable.
The first study covers 50 customer-support messages, Jev, and one comparison model. Accuracy/agreement, latency, and cost are planned measurements, not recorded results.
- 01
Define the task
Begin with a fixed set of 50 customer-support messages and record the task framing before any comparison is run.
- 02
Set the boundaries
Resolve how greetings, out-of-scope messages, and multi-intent requests will be treated before labeling. A clean label set starts with explicit edges.
- 03
Run the comparison
Compare Jev with one comparison model against the final protocol. The study remains planned until it is actually executed.
- 04
Publish the ledger
Report classification agreement, response latency, and cost only after measurement, with the method beside each result.
02 / study card
The planned scope.
Inputs, labels, and outputs are defined here before the first run.
Before the first run
- Which out-of-scope greetings should be excluded?
- How should multi-intent messages be labeled?
- Is a five-label forced classification protocol valid for this sample?
DEC-001 / planned study
The method comes before the score.
The first entry will publish classification agreement, response latency, and cost only after the sample and label rules are settled.
Read how this static site works ↗