Demos
Interactive walkthroughs of what gutcheck does. Demos marked illustrative use hand-picked values to show behaviour. The measured results section uses real numbers from the committed evals.
Decision playground
Ask typed questions, then move the thresholds and watch verdicts change. The probabilities stay fixed, which is the point: the model's belief and your risk appetite are separate things.
We were billed twice for invoice #4411 and I want the extra charge back.
departmentactWhich department should handle this?
Safe to act on automatically.
refundreviewDoes the customer ask for money back?
Queue for a quick human or rule check.
urgentreviewDoes the customer need help within the hour?
Queue for a quick human or rule check.
Drag the thresholds. The model's probabilities stay put; only the verdicts move. In the API you set these globally, per request or per question.
{
"state": "We were billed twice for invoice #4411 and I want the extra charge back.",
"policy": {
"act_at": 0.9,
"review_at": 0.6
},
"questions": {
"department": {
"type": "choice",
"instructions": "Which department should handle this?"
},
"refund": {
"type": "noul",
"instructions": "Does the customer ask for money back?"
},
"urgent": {
"type": "noul",
"instructions": "Does the customer need help within the hour?"
}
}
}{
"answers": {
"department": {
"type": "choice",
"answer": "billing",
"answer_probability": 0.91,
"verdict": "act"
},
"refund": {
"type": "noul",
"answer": "yes",
"answer_probability": 0.82,
"verdict": "review"
},
"urgent": {
"type": "noul",
"answer": "no",
"answer_probability": 0.69,
"verdict": "review"
}
}
}Illustrative responses in the real response format. Probabilities are hand-picked to show behaviour, not live model output. For measured numbers, see the eval results.
Calibration lab
A model can be right and still sound too sure. Temperature scaling fixes the sound without changing the answer.
The encoder returned logits [3.2, 1.1, 0.4]. Divide by T, then softmax. The winner never changes; the confidence does.
Verdict thresholds: act at 0.90, review at 0.60. T above 1 softens an overconfident model; T below 1 sharpens an underconfident one. gutcheck fits one T per question by minimising negative log-likelihood on labelled data.
Of the answers given at 90-100% confidence, how many were right? Illustrative curves for an overconfident model.
confidence claimed accuracy achieved. When the bars match, a 0.9 really means 9 in 10.
The feedback loop
Your app asks a question. gutcheck answers and logs the raw probabilities under a trace id.
curl localhost:8080/v1/decide -d '{
"state": "Please cancel my order",
"questions": {"urgent": {"type": "noul",
"instructions": "Is this urgent?"}}
}'
{ "trace_id": "gc_4f9c...",
"answers": {"urgent": {"noul": 0.81, "verdict": "review"}} }Response values are illustrative.
Question packs
id: support-triage
version: 1
description: Routes support tickets.
questions:
urgent:
type: noul
instructions: Does the customer need help within the hour?
policy: {act_at: 0.95, review_at: 0.7}
eval:
test: {path: test.jsonl, license: CC-BY-4.0}
calibration: {path: train.jsonl, license: CC-BY-4.0}
text_field: text
label_field: label
labels: {"urgent": true, "normal": false}A pack is a directory: questions, the datasets that test them, a fitted calibration and a generated report. The numbers in the eval output above are an example, not a real pack.
Measured results
The bundled prompt-guard pack, base model against the fine-tuned checkpoint. Fine-tuning is what moves most answers from "escalate" to "act".
higher is better
prompt-guard.injectionprompt-guard.jailbreakMeasured on held-out test splits the model never saw: deepset/prompt-injections (116 rows) and jackhhao/jailbreak-classification (400 rows). Source: the committed EVAL.md. Bars for calibration error are scaled for visibility.
More detail, including the exact datasets and revisions, is on the prompt-guard page.
Run the real thing
docker build -t gutcheck . && docker run -p 8080:8080 -v gutcheck-data:/data gutcheck
curl localhost:8080/v1/packsThen open http://localhost:8080/dashboard. Full setup is in the quickstart.