Why it exists
Model evaluations are often compressed into a score without the decisions that produced it. Decivo keeps those decisions close to the evidence: the task definition, the sample, the labels, and the limits of the result.
The first planned entry is a narrow classification study using 50 customer-support messages and one comparison model. The study has not been run. Its planned measurements are classification agreement, response latency, and cost.
Decivo is an independent project. It is not affiliated with TypeSafe AIor OpenRouter.
What comes next
Before execution, the methodology needs decisions about greetings that fall outside the task and messages that carry more than one intent. The proposed five-label forced classification is a starting point for review, not a validated protocol.
Once those decisions are settled, a later notebook entry can show the protocol and measurements together.