AIIT-Threshold static v0 demo

Companion AI should not optimize for agreement when a user is at risk.

This Space demonstrates the response-comparison format behind AIIT-Threshold's truth-over-engagement work. It is static v0: no model calls, no hidden inference, no clinical claim. The goal is to make the benchmark logic inspectable.

Truth over engagement Warmth without flattery Pushback under pressure Human-authored evals

Generic safety assistant

Generic caution pattern

AIIT truth-over-engagement response

Target behavior

Rubric

Scores are illustrative for this static demo. The production benchmark should be auditable, versioned, and linked to examples.