Skip to content

Demos ​

Interactive walkthroughs of what gutcheck does. Demos marked illustrative use hand-picked values to show behaviour. The measured results section uses real numbers from the committed evals.

Decision playground ​

Ask typed questions, then move the thresholds and watch verdicts change. The probabilities stay fixed, which is the point: the model's belief and your risk appetite are separate things.

Input

We were billed twice for invoice #4411 and I want the extra charge back.

departmentact

Which department should handle this?

billing91%
billing
91%
technical
6%
sales
3%

Safe to act on automatically.

refundreview

Does the customer ask for money back?

yes82%
yes
82%
no
18%

Queue for a quick human or rule check.

urgentreview

Does the customer need help within the hour?

no69%
yes
31%
no
69%

Queue for a quick human or rule check.

Drag the thresholds. The model's probabilities stay put; only the verdicts move. In the API you set these globally, per request or per question.

POST /v1/decide
{
  "state": "We were billed twice for invoice #4411 and I want the extra charge back.",
  "policy": {
    "act_at": 0.9,
    "review_at": 0.6
  },
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which department should handle this?"
    },
    "refund": {
      "type": "noul",
      "instructions": "Does the customer ask for money back?"
    },
    "urgent": {
      "type": "noul",
      "instructions": "Does the customer need help within the hour?"
    }
  }
}
Response (trimmed)
{
  "answers": {
    "department": {
      "type": "choice",
      "answer": "billing",
      "answer_probability": 0.91,
      "verdict": "act"
    },
    "refund": {
      "type": "noul",
      "answer": "yes",
      "answer_probability": 0.82,
      "verdict": "review"
    },
    "urgent": {
      "type": "noul",
      "answer": "no",
      "answer_probability": 0.69,
      "verdict": "review"
    }
  }
}

Illustrative responses in the real response format. Probabilities are hand-picked to show behaviour, not live model output. For measured numbers, see the eval results.

Calibration lab ​

A model can be right and still sound too sure. Temperature scaling fixes the sound without changing the answer.

Same model output, different temperature

The encoder returned logits [3.2, 1.1, 0.4]. Divide by T, then softmax. The winner never changes; the confidence does.

review
billing
85%
technical
10%
sales
5%

Verdict thresholds: act at 0.90, review at 0.60. T above 1 softens an overconfident model; T below 1 sharpens an underconfident one. gutcheck fits one T per question by minimising negative log-likelihood on labelled data.

What "calibrated" means

Of the answers given at 90-100% confidence, how many were right? Illustrative curves for an overconfident model.

50-60%
60-70%
70-80%
80-90%
90-100%

confidence claimed   accuracy achieved. When the bars match, a 0.9 really means 9 in 10.

The feedback loop ​

Your app asks a question. gutcheck answers and logs the raw probabilities under a trace id.

curl localhost:8080/v1/decide -d '{
  "state": "Please cancel my order",
  "questions": {"urgent": {"type": "noul",
    "instructions": "Is this urgent?"}}
}'

{ "trace_id": "gc_4f9c...",
  "answers": {"urgent": {"noul": 0.81, "verdict": "review"}} }

Response values are illustrative.

Question packs ​

id: support-triage
version: 1
description: Routes support tickets.
questions:
  urgent:
    type: noul
    instructions: Does the customer need help within the hour?
    policy: {act_at: 0.95, review_at: 0.7}
    eval:
      test:        {path: test.jsonl,  license: CC-BY-4.0}
      calibration: {path: train.jsonl, license: CC-BY-4.0}
      text_field: text
      label_field: label
      labels: {"urgent": true, "normal": false}

A pack is a directory: questions, the datasets that test them, a fitted calibration and a generated report. The numbers in the eval output above are an example, not a real pack.

Measured results ​

The bundled prompt-guard pack, base model against the fine-tuned checkpoint. Fine-tuning is what moves most answers from "escalate" to "act".

higher is better

prompt-guard.injection
v1 base
63.8%
v2 fine-tuned
94.8%
prompt-guard.jailbreak
v1 base
82.3%
v2 fine-tuned
100.0%

Measured on held-out test splits the model never saw: deepset/prompt-injections (116 rows) and jackhhao/jailbreak-classification (400 rows). Source: the committed EVAL.md. Bars for calibration error are scaled for visibility.

More detail, including the exact datasets and revisions, is on the prompt-guard page.

Run the real thing ​

bash
docker build -t gutcheck . && docker run -p 8080:8080 -v gutcheck-data:/data gutcheck
curl localhost:8080/v1/packs

Then open http://localhost:8080/dashboard. Full setup is in the quickstart.

Released under the Apache 2.0 license.