Independent research notebook / vol. 01

Jev classification, tested in the open.

Decivo documents a narrow benchmark of Jev, TypeSafe AI's typed decision model, on a fixed customer-support sample. The study is in preparation; results are pending.

Protocol
Planned / not run
Sample
50 messages / proposed
Format
Static notebook / decivo.org

00 / premise

A narrow benchmark for typed decisions.

Jev is designed to return typed decisions from unstructured input. Decivo asks one bounded question: how does it classify 50 customer-support messages against one comparison model when the labels and edge cases are stated up front?

01 / method

Make each decision inspectable.

The first study covers 50 customer-support messages, Jev, and one comparison model. Accuracy/agreement, latency, and cost are planned measurements, not recorded results.

  1. 01

    Define the task

    Begin with a fixed set of 50 customer-support messages and record the task framing before any comparison is run.

  2. 02

    Set the boundaries

    Resolve how greetings, out-of-scope messages, and multi-intent requests will be treated before labeling. A clean label set starts with explicit edges.

  3. 03

    Run the comparison

    Compare Jev with one comparison model against the final protocol. The study remains planned until it is actually executed.

  4. 04

    Publish the ledger

    Report classification agreement, response latency, and cost only after measurement, with the method beside each result.

02 / study card

The planned scope.

Inputs, labels, and outputs are defined here before the first run.

DEC-001 / study briefstate: Results pending
Sample50customer-support messages
Comparison1model planned
Outputs3metrics proposed
Accuracy / agreementplanned measurement
Latencyplanned measurement
Costplanned measurement

Before the first run

  • Which out-of-scope greetings should be excluded?
  • How should multi-intent messages be labeled?
  • Is a five-label forced classification protocol valid for this sample?

DEC-001 / planned study

The method comes before the score.

The first entry will publish classification agreement, response latency, and cost only after the sample and label rules are settled.

Read how this static site works ↗