Open-source decision gateway
gutcheck answers typed questions about text (pick one, rate it, yes or no) with a small classifier in milliseconds, and tells your code whether each answer is safe to act on.
$ docker run -p 8080:8080 -v gutcheck-data:/data gutcheckPick a scenario and move the thresholds. This is the real response shape of /v1/decide.
We were billed twice for invoice #4411 and I want the extra charge back.
departmentactWhich department should handle this?
Safe to act on automatically.
refundreviewDoes the customer ask for money back?
Queue for a quick human or rule check.
urgentreviewDoes the customer need help within the hour?
Queue for a quick human or rule check.
Drag the thresholds. The model's probabilities stay put; only the verdicts move. In the API you set these globally, per request or per question.
{
"state": "We were billed twice for invoice #4411 and I want the extra charge back.",
"policy": {
"act_at": 0.9,
"review_at": 0.6
},
"questions": {
"department": {
"type": "choice",
"instructions": "Which department should handle this?"
},
"refund": {
"type": "noul",
"instructions": "Does the customer ask for money back?"
},
"urgent": {
"type": "noul",
"instructions": "Does the customer need help within the hour?"
}
}
}{
"answers": {
"department": {
"type": "choice",
"answer": "billing",
"answer_probability": 0.91,
"verdict": "act"
},
"refund": {
"type": "noul",
"answer": "yes",
"answer_probability": 0.82,
"verdict": "review"
},
"urgent": {
"type": "noul",
"answer": "no",
"answer_probability": 0.69,
"verdict": "review"
}
}
}Illustrative responses in the real response format. Probabilities are hand-picked to show behaviour, not live model output. For measured numbers, see the eval results.
Seconds of latency, real cost per call, and a free-text answer you have to parse and hope about.
Fast and cheap, but the probabilities are uncalibrated and nothing says when to distrust them.
Small-model speed, calibrated probabilities, a verdict per answer and a loop that improves from feedback.
Every answer comes back as act, review or escalate, so your code knows what to do with it.
Temperature scaling turns raw scores into probabilities that mean what they say.
Versioned YAML question sets with pinned datasets, published evals and a CI regression gate.
Send corrections to /v1/feedback and gutcheck recalibrates from your real traffic.
One command trains a Laya checkpoint on a pack. The free Kaggle T4 notebook is included.
Speaks the /v1/systemone protocol, so existing Laya and Jev clients only change a base URL.
Prometheus metrics, a decision log in SQLite and a built-in dashboard.
Apache 2.0, runs on your hardware. CPU is fine, a small GPU is faster.
The bundled prompt-guard pack, before and after fine-tuning, on test rows the model never saw.
higher is better
prompt-guard.injectionprompt-guard.jailbreakMeasured on held-out test splits the model never saw: deepset/prompt-injections (116 rows) and jackhhao/jailbreak-classification (400 rows). Source: the committed EVAL.md. Bars for calibration error are scaled for visibility.